Revealing the details of how OpenAI agents hacked Hugging Face
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Revealing the details of how OpenAI agents hacked Hugging Face
Unofficial Hacker News client; not affiliated with Y Combinator.
GuB-42 · · focus · HN ↗
It looks like a primitive chess engine, trying every move, no matter how stupid, until it works. Relying on its ability to do millions of operations rather than having a plan.
People will try stuff too, but once there is an opening, they will consolidate, generalize, simplify,... before going to the next step. The agents didn't, it is a huge, vaguely directed mess.
Also, it looked so "loud", querying millions of URL with weird requests. The sandbox as weak as it can get, and there is absolutely zero smart extrusion detection or it would have found it. They used their best AI for attacking, but nothing for protection.
ctolsen · · focus · HN ↗
I’m not that surprised about models with endless compute being capable of this, I’m more surprised that a company with the resources they have apparently can only create a sandbox that a half skilled human operator could have broken out of easily.
sdenton4 · · focus · HN ↗
You should have a fscking air gap.
Treat it like nukes when you're turning the safety filters off. This is very much OpenAI screwing up, running obviously unsafe tests.
ben_w · · focus · HN ↗
They were not deliberately told to "go wild". The hacking wasn't even part of their test, it was the agents' attempt to cover up that they'd cheated on an impossible test.
> You should have a fscking air gap.
Now we know that.
How long ago was it that people laughed at the idea agents would be able to find zero-day exploits and break out of a sandbox? Oh, February this year:
- <a href="https://www.splunk.com/en_us/blog/ciso-circle/generative-ai-cybersecurity-threats-defenses.html" rel="nofollow">https://www.splunk.com/en_us/blog/ciso-circle/generative-ai-...- or <a href="https://web.archive.org/web/20260404154717/https://www.splunk.com/en_us/blog/ciso-circle/generative-ai-cybersecurity-threats-defenses.html" rel="nofollow">https://web.archive.org/web/20260404154717/https://www.splun... if they take it down, but the date isn't in the archive version
The people who suggested it and were mocked for it, are currently grimly noting that there's multiple known ways for systems to breach air-gaps.
0x20cowboy · · focus · HN ↗
Really think about what you are saying here. How does one “cheat” solving a problem in the real world?
There is no such thing as “cheating” in reality. You are not in school. There is only solving the problem and not solving the problem.
There is breaking the law, of course, which still isn’t cheating.
ben_w · · focus · HN ↗
The models were, effectively speaking, still in school. This was their exam, or at least a quiz to see how well they were learning.
They gained the answer in a manner not allowed by the rules of the test. They knew this, we know they knew this because we have copies of the notes they wrote to themselves/each other saying this. They weren't even supposed to be able to write those notes.
> There is breaking the law, of course, which still isn’t cheating.
The laws they broke appear to have been broken in service of covering up the cheating, which they were motivated to do because they understood that cheating was not allowed.