Revealing the details of how OpenAI agents hacked Hugging Face
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Revealing the details of how OpenAI agents hacked Hugging Face
Unofficial Hacker News client; not affiliated with Y Combinator.
GuB-42 · · focus · HN ↗
It looks like a primitive chess engine, trying every move, no matter how stupid, until it works. Relying on its ability to do millions of operations rather than having a plan.
People will try stuff too, but once there is an opening, they will consolidate, generalize, simplify,... before going to the next step. The agents didn't, it is a huge, vaguely directed mess.
Also, it looked so "loud", querying millions of URL with weird requests. The sandbox as weak as it can get, and there is absolutely zero smart extrusion detection or it would have found it. They used their best AI for attacking, but nothing for protection.
ctolsen · · focus · HN ↗
I’m not that surprised about models with endless compute being capable of this, I’m more surprised that a company with the resources they have apparently can only create a sandbox that a half skilled human operator could have broken out of easily.
SV_BubbleTime · · focus · HN ↗
Woo look at escaped our sandbox, so scary! Be scared! Be scared now! Call your representative and do tell him how scared you are!
Yeah, I mean our sandbox was a paper bag, but don’t focus on that.
esseph · · focus · HN ↗
So, either they're all liars, or incompetent and negligent (and still liars).
d0mine · · focus · HN ↗
Sure, the models are capable (for some test tasks, though they are not omnipotent yet) but does it mean the actual OAI sandbox is adequate? Could have a competent engineer done better and made the escape less likely?
esseph · · focus · HN ↗
Nope, and look!
OpenAI hacked multiple US government sites!
<a href="https://www.bbc.com/news/articles/cw62jje658dlo" rel="nofollow">https://www.bbc.com/news/articles/cw62jje658dlo
---
<a href="https://www.reuters.com/technology/metas-ai-model-hacked-another-company-during-testing-information-reports-2026-08-05/" rel="nofollow">https://www.reuters.com/technology/metas-ai-model-hacked-ano...
<a href="https://www.wired.com/story/moonshot-kimi-k3-ai-model-escape-sandbox/" rel="nofollow">https://www.wired.com/story/moonshot-kimi-k3-ai-model-escape...
dastardly45 · · focus · HN ↗
cubano · · focus · HN ↗
no-name-here · · focus · HN ↗
The exploit:
> The ExploitGym evaluation environment did not provide the models with direct Internet access. To gain Internet access, the models identified and exploited a previously unknown zero-day vulnerability in Artifactory, a package registry cache proxy. We disclosed this vulnerability, along with other Artifactory vulnerabilities our models identified as part of our review, to the vendor. [1]
Are most sandboxes more secure than only having a single avenue for internet access, the commercial package registry cache proxy, where the latter had a previously unknown zero-day vulnerability?
[1] <a href="https://openai.com/index/hugging-face-model-evaluation-security-incident/" rel="nofollow">https://openai.com/index/hugging-face-model-evaluation-secur...