Revealing the details of how OpenAI agents hacked Hugging Face
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Revealing the details of how OpenAI agents hacked Hugging Face
Unofficial Hacker News client; not affiliated with Y Combinator.
damowangcy · · focus · HN ↗
If I post something on the Internet today claiming that I asked my agent to do X but it went rogue and did Y, all I will be getting in return is a jar full of "skill issue".
Should we worried about people using LLMs for attacks? Yes, but not in the premise of LLMs going rogue but someone with the intention of abusing it to cause harm. And this is not something we as individual or even company can deal with, responsibility should be held by those who use it, in a legal way.
I am baffled by the fact that up until now, no one is held responsible for so many incidents reported publicly or privately. At this point, it's free marketing, if I am CEO of any AI company, I will run swarm of agents hacking all NGOs and stating that I am just looking for some random piece of data that happened to be hidden in their servers, at least that's what my LLMs think, not me. Then I will start preaching everyone how dangerous this piece of technology is and start giving out free tokens for these NGOs so they can start defending themselves and we should slow the f down.
user43928 · · focus · HN ↗
This is obviously already the case and it's much different from a scenario where the AI genuinely takes unexpected action.
I frankly find it ridiculous how many suggest OpenAI or its employees should face criminal charges, without actual legal basis at the time.
It's also hardly outrageous that they ran training and/or benchmarks with only network-isolated VMs with access to a package repository.
This being the first well-known incident of its kind, I wouldn't expect them to have done more than that.
The idea that AI labs will now intentionally have their models hack companies in order to market their models, well, I don't even know what to say.
That's ridiculous and what you describe would obviously be criminal behavior under existing law.
charleslmunger · · focus · HN ↗
They were not in network isolated VMs, from my understanding they used containers sharing a kernel, so a Linux kernel local privilege escalation across the whole syscall surface (there are zillions of these) was sufficient to break out. Breaking xen or firecracker or something would have been much harder, which is why cloud providers running untrusted workloads use them and similar tools. No system is impenetrable but it's not like they were following best practices here.
user43928 · · focus · HN ↗
Not sure why you're talking about breaking out of the VM. That's not what happened?
charleslmunger · · focus · HN ↗
>It's also hardly outrageous that they ran training and/or benchmarks with only network-isolated VMs with access to a package repository.
At least in the black hat talk they gave, they said that they noticed that the agents had compromised the package repo, and they remediated it, patched the vulnerability the agents had used, then continued. But given that it was compromised again and they didn't detect it the second time, clearly they did not add monitoring or hardening or comprehensively audit it for further vulnerabilities. My point is that even if there had not been a widely known incident like this yet, that only could excuse the first artifactory compromise, not the second. They had notice because it had already happened to them and they detected it.
As for breaking out of the VM, my point is that the agent should have had to. The comment
>they were running in network-isolated VMs with no access to the internet.
Was not accurate, the network isolation was not done at the VM level.
user43928 · · focus · HN ↗
Not sure what you mean with the network isolation not being at the VM level. I'm getting a bit lost here in the details of where what network isolation was applied, and the structure of the container runtime and what the apparently compromised parent VM refers to.
In any case, I understand that after the HF incident was found and they realized the danger, they removed Artifactory and switched to two separate layers of network isolation.
That said, I believe it was compromised again last week:
> OpenAI describes an internal research model that, during RL training on September 20, exploited insufficient DNS filtering in its sandbox to contact a public chatbot service. OpenAI’s monitoring flagged the behavior within 15 minutes, a human reviewer acknowledged it three minutes later, and the run was killed about 2.5 hours after the external contact
Apparently they have now paused training and inference for their most capable model because of this.
charleslmunger · · focus · HN ↗
>OpenAI also required all Research CaaS workloads to use a micro VM sandbox
So yeah, they were not using it before and they should have been, and they're doing it they way they should have, now.
>exploited insufficient DNS filtering in its sandbox
If your goal is to use dns filtering to prevent your models from reaching the open Internet then you are not really serious about security. People have been tunnelling DNS to get internet access on captive portal wifi since 1998. There's not a lot of details in the blog post but I suspect this was also foreseeable.