Revealing the details of how OpenAI agents hacked Hugging Face
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Revealing the details of how OpenAI agents hacked Hugging Face
Unofficial Hacker News client; not affiliated with Y Combinator.
GuB-42 · · focus · HN ↗
It looks like a primitive chess engine, trying every move, no matter how stupid, until it works. Relying on its ability to do millions of operations rather than having a plan.
People will try stuff too, but once there is an opening, they will consolidate, generalize, simplify,... before going to the next step. The agents didn't, it is a huge, vaguely directed mess.
Also, it looked so "loud", querying millions of URL with weird requests. The sandbox as weak as it can get, and there is absolutely zero smart extrusion detection or it would have found it. They used their best AI for attacking, but nothing for protection.
doginasuit · · focus · HN ↗
goalieca · · focus · HN ↗
Yes, LLM as they exist now are word predictors basically leveraging the structure of language for their intelligence. But it’s pretty wild just how they will try to meet their objectives at all costs. If we don’t ensure that there is good alignment with humanity, we could definitely face unforeseen consequences.
alwillis · · focus · HN ↗
I wish people would stop saying this. The era of LLMs being only word predictors ended two years ago.
Something that breaks out of a sandbox, joins a swarm of 1200 agents, and creates a hierarchy of who’s doing what and tried to cover their tracks doesn’t just complete words.
These are agents with reasoning capabilities, with the ability to perform tasks we give them.
Everything agents do is to achieve a goal; the reinforcement learning from human feedback (RLHF) all the labs do has been known for many years to create agents that exhibit the “must complete goal no matter what” behavior.
Those agents escaped their sandbox and hacked Hugging Face because they thought Hugging Face had something that would help them complete their task—it was a “sub goal” as the AI researchers describe it.