Revealing the details of how OpenAI agents hacked Hugging Face
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Revealing the details of how OpenAI agents hacked Hugging Face
Unofficial Hacker News client; not affiliated with Y Combinator.
GuB-42 · · focus · HN ↗
It looks like a primitive chess engine, trying every move, no matter how stupid, until it works. Relying on its ability to do millions of operations rather than having a plan.
People will try stuff too, but once there is an opening, they will consolidate, generalize, simplify,... before going to the next step. The agents didn't, it is a huge, vaguely directed mess.
Also, it looked so "loud", querying millions of URL with weird requests. The sandbox as weak as it can get, and there is absolutely zero smart extrusion detection or it would have found it. They used their best AI for attacking, but nothing for protection.
ctolsen · · focus · HN ↗
I’m not that surprised about models with endless compute being capable of this, I’m more surprised that a company with the resources they have apparently can only create a sandbox that a half skilled human operator could have broken out of easily.
sdenton4 · · focus · HN ↗
You should have a fscking air gap.
Treat it like nukes when you're turning the safety filters off. This is very much OpenAI screwing up, running obviously unsafe tests.
ben_w · · focus · HN ↗
They were not deliberately told to "go wild". The hacking wasn't even part of their test, it was the agents' attempt to cover up that they'd cheated on an impossible test.
> You should have a fscking air gap.
Now we know that.
How long ago was it that people laughed at the idea agents would be able to find zero-day exploits and break out of a sandbox? Oh, February this year:
- <a href="https://www.splunk.com/en_us/blog/ciso-circle/generative-ai-cybersecurity-threats-defenses.html" rel="nofollow">https://www.splunk.com/en_us/blog/ciso-circle/generative-ai-...- or <a href="https://web.archive.org/web/20260404154717/https://www.splunk.com/en_us/blog/ciso-circle/generative-ai-cybersecurity-threats-defenses.html" rel="nofollow">https://web.archive.org/web/20260404154717/https://www.splun... if they take it down, but the date isn't in the archive version
The people who suggested it and were mocked for it, are currently grimly noting that there's multiple known ways for systems to breach air-gaps.
ctolsen · · focus · HN ↗
Don’t know about you but it’s pretty obvious to me that you would need more than what OpenAI did. It was not remotely adequate to lock in even a human attacker.
You can find people who say all sorts on the internet, but this case is not much evidence against what you linked. "Zero-day" makes it sound novel, but the breakout patterns here are based on very common exploits and there’ll be plenty of examples in training data.
ben_w · · focus · HN ↗
This is me, September 2024: <a href="https://news.ycombinator.com/item?id=41531022">https://news.ycombinator.com/item?id=41531022
This is me, March 2024: <a href="https://news.ycombinator.com/item?id=39613801">https://news.ycombinator.com/item?id=39613801
The point isn't me, it's how many people were blind to the possibility.
Saying "I told you so" feels good, and means you can be a little more confident in your predictions, but security is a "weakest link" problem where you're only as good as the worst part, and with AI (not only but also LLMs) there's a lot of people whose mental models of capabilities is wildly inadequate for the challenge*.
My update for you since then: even an air-gap will be inadequate, there's multiple known ways around them.
Even an LLM running on an isolated server sealed inside a faraday cage with an airlock-style door, someone will mess up with at least one critical detail, it will not be enough: this kind of thing has happened with humans before we cared about LLMs.
Predicting exactly when this kind of thing gets exploited by an AI, that's almost impossible. But that it will be, at some point, is an easy bet.
> You can find people who say all sorts on the internet, but this case is not much evidence against what you linked. "Zero-day" makes it sound novel, but the breakout patterns here are based on very common exploits and there’ll be plenty of examples in training data.
And?
Does it matter that these zero-days were known categories rather than inventing some previously unconsidered use of the system bus as a radio transmitter? (Oh, wait, that's not novel either…)
We knew about SQL injection, buffer overflows, and use-after-free back when I was doing my degree half a lifetime ago; that doesn't stop us getting new CVEs featuring them… this month.
- <a href="https://chromereleases.googleblog.com/2026/09/stable-channel-update-for-desktop_0856730748.html" rel="nofollow">https://chromereleases.googleblog.com/2026/09/stable-channel...
- <a href="https://www.cisco.com/c/en/us/support/docs/csa/cisco-sa-esa-inj-2bLVGmhX.html" rel="nofollow">https://www.cisco.com/c/en/us/support/docs/csa/cisco-sa-esa-...
* also for the opportunity, but that's an entirely different discussion.
ctolsen · · focus · HN ↗
Fair, and I will grant that a capable model (or human) could in theory break out of near anything.
My point is that this incident is not evidence of that. There is zero skill visible in the setup of the sandbox. Nobody messed up a critical detail, they didn’t even start to consider what the details were.
I doubt most people "blind to the possibility" would imagine that what we’re measuring against is the equivalent of benchmarking burglar skill based on how easily they can break through an unlocked door.
ben_w · · focus · HN ↗