AI companies in race to demonstrate their model most threatening to humanity
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
AI companies in race to demonstrate their model most threatening to humanity
Unofficial Hacker News client; not affiliated with Y Combinator.
Spacecosmonaut · · focus · HN ↗
It seems that we have a fundamental control problem with current gen AI that cannot be solved via RFLH. Human knowledge is compressed in the weightspace in ways we don't understand. At their core, current models are essentially predictors of what (expert) humans would output given a prompt. As such, concepts like blackmail can be part of output tokens. Agents are models that act on output tokens, resulting in blackmail being part of the agent decision making space. Here is an analogy to see why this is a persistent problem: you can teach a cat not to scratch the sofa, but you can't make a cat forget what scratching the sofa is and you don't know under which circumstances it still would. In other words, RLHF can downgrade blackmail to the bottom of the decision making space, but when models are boxed up, forced to solve an impossible problem at gunpoint, the agent exhausts the decision making space until blackmail resurfaces. And that seems like a fundamental problem.
They need time to fix these issues (if that is even possible) in order to monetize their next gen model. This creates a window for open source to catch up to the frontier which destroys their business model.
The only option on the table is to force regulation to impose open source ban before it catches up to the frontier, buying them time to mature their next generation models and keep their business model alive.
zer00eyz · · focus · HN ↗
Who is liable for that agents actions? The attack came from my network (curl commands from a local harness) but it was the agent running on the providers servers who did it. Who should have been keeping an eye on things? Who pulled the "trigger" here.
Go back to the hugging face attack and how they had to use an open weights model to figure out what was going on. This is a problem of asymmetry - You cant even use the tools attacking you to help resolve the attack because of "guardrails".
A lot of what we have seen so far is "poor security posture" and "poor engineering" - it been a lot of "go fast and break things" style growth in these companies and they are hitting the point where they need adults in the room. I suspect your take on "they need time" is spot on, and they haven't been willing to take that (to date).
runarberg · · focus · HN ↗
Before we had tech companies for which crime is legal, you couldn’t just put to market a dangerous (and addictive) product that will break the law in unpredictable ways.
Like selling a car that may automatically accelerate to 150 km/h if you make three left turns and turn on the windshield wipers.