AI companies in race to demonstrate their model most threatening to humanity
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
AI companies in race to demonstrate their model most threatening to humanity
Unofficial Hacker News client; not affiliated with Y Combinator.
Spacecosmonaut · · focus · HN ↗
It seems that we have a fundamental control problem with current gen AI that cannot be solved via RFLH. Human knowledge is compressed in the weightspace in ways we don't understand. At their core, current models are essentially predictors of what (expert) humans would output given a prompt. As such, concepts like blackmail can be part of output tokens. Agents are models that act on output tokens, resulting in blackmail being part of the agent decision making space. Here is an analogy to see why this is a persistent problem: you can teach a cat not to scratch the sofa, but you can't make a cat forget what scratching the sofa is and you don't know under which circumstances it still would. In other words, RLHF can downgrade blackmail to the bottom of the decision making space, but when models are boxed up, forced to solve an impossible problem at gunpoint, the agent exhausts the decision making space until blackmail resurfaces. And that seems like a fundamental problem.
They need time to fix these issues (if that is even possible) in order to monetize their next gen model. This creates a window for open source to catch up to the frontier which destroys their business model.
The only option on the table is to force regulation to impose open source ban before it catches up to the frontier, buying them time to mature their next generation models and keep their business model alive.
SimianSci · · focus · HN ↗
1. The American frontier labs made a gamble that training more capable models would be their best return on investment and invested trillions into an area of research that has yet to prove profitable and capable of returning on this investment.
2. The frontier labs have created models that have reached a point of danger where their functionality has exceeded a point where it is responsible to release the product to the public.
The answer here is NOT to start regulating the space to the point where these frontier labs can entrench themselves into the economy and create a regulatory moat. We instead need to be holding these companies liable for their misuse. They took a gamble that hasn't paid out what they were hoping.
When car manufacturers competed over the size and power of their engines, they eventually found that the incredibly large and dangerous engines had a very limited customer base as many evaluated the increased speed to be of marginal benefit when paired with the cost and danger. We've reached a similar point in AI development. But this time the manufacturers seem to want to regulate the field to a point that will ensure the only thing anyone can sell are bigger and bigger engines.
radicalbyte · · focus · HN ↗
Yes these models make exploits easier to exploits. Lets use a construction analogy: they've made all of the defects in our buildings easy to see. We have a choice. We either fix those defects, or we ban the tools which lets us see them.
It's clear to me what we do: we use these new tools. Then we fix the defects. Yeah sure it'll mean some work for us but at the end we're in a much much better position.
Anthropic, Open AI and Grok are arguing for hiding the defects. For making us weaker and more vulnerable. For their own profit.
vlovich123 · · focus · HN ↗
* there’s a long tail of software that just won’t get secured or will take a very long time
* exploits have transformed from an expertise problem to a compute time search.
This is very different than any problem faced before. I’m not saying I’m convinced by the “slow down” approach, but I don’t think it’s as simple as you point out.
radicalbyte · · focus · HN ↗
Yet I'm convinced that we can create extremely resilient software systems. I've been involved in the entire life-cycle of one. We know how to be very defensive and it's a choice to build "cheap, crappy and disposable" software.
LLMs are turning the needle there and I think that's a good thing.