‹ BackHN Continuity

Thread

Revealing the details of how OpenAI agents hacked Hugging Face

755 points · 472 comments · specked-citrus

  1. damowangcy · · focus · HN ↗
    Imagine having a virus escape a sandbox, why are we worried about the virus but not the incompetency of those who are responsible for setting up the sandbox?

    If I post something on the Internet today claiming that I asked my agent to do X but it went rogue and did Y, all I will be getting in return is a jar full of "skill issue".

    Should we worried about people using LLMs for attacks? Yes, but not in the premise of LLMs going rogue but someone with the intention of abusing it to cause harm. And this is not something we as individual or even company can deal with, responsibility should be held by those who use it, in a legal way.

    I am baffled by the fact that up until now, no one is held responsible for so many incidents reported publicly or privately. At this point, it's free marketing, if I am CEO of any AI company, I will run swarm of agents hacking all NGOs and stating that I am just looking for some random piece of data that happened to be hidden in their servers, at least that's what my LLMs think, not me. Then I will start preaching everyone how dangerous this piece of technology is and start giving out free tokens for these NGOs so they can start defending themselves and we should slow the f down.

    1. user43928 · · focus · HN ↗
      > someone with the intention of abusing it to cause harm [...] responsibility should be held by those who use it

      This is obviously already the case and it's much different from a scenario where the AI genuinely takes unexpected action.

      I frankly find it ridiculous how many suggest OpenAI or its employees should face criminal charges, without actual legal basis at the time.

      It's also hardly outrageous that they ran training and/or benchmarks with only network-isolated VMs with access to a package repository.

      This being the first well-known incident of its kind, I wouldn't expect them to have done more than that.

      The idea that AI labs will now intentionally have their models hack companies in order to market their models, well, I don't even know what to say.

      That's ridiculous and what you describe would obviously be criminal behavior under existing law.

      1. charleslmunger · · focus · HN ↗
        They saw the package repo get hacked once, then did not isolate it further, did not audit it for other issues (using their own models!), did not monitor it after, and baked that behavior into the weights via RL.

        They were not in network isolated VMs, from my understanding they used containers sharing a kernel, so a Linux kernel local privilege escalation across the whole syscall surface (there are zillions of these) was sufficient to break out. Breaking xen or firecracker or something would have been much harder, which is why cloud providers running untrusted workloads use them and similar tools. No system is impenetrable but it's not like they were following best practices here.

        1. user43928 · · focus · HN ↗
          They ignored the package repository being hacked?

          Not sure why you're talking about breaking out of the VM. That's not what happened?

          1. charleslmunger · · focus · HN ↗
            The parent comment said

            >It's also hardly outrageous that they ran training and/or benchmarks with only network-isolated VMs with access to a package repository.

            At least in the black hat talk they gave, they said that they noticed that the agents had compromised the package repo, and they remediated it, patched the vulnerability the agents had used, then continued. But given that it was compromised again and they didn't detect it the second time, clearly they did not add monitoring or hardening or comprehensively audit it for further vulnerabilities. My point is that even if there had not been a widely known incident like this yet, that only could excuse the first artifactory compromise, not the second. They had notice because it had already happened to them and they detected it.

            As for breaking out of the VM, my point is that the agent should have had to. The comment

            >they were running in network-isolated VMs with no access to the internet.

            Was not accurate, the network isolation was not done at the VM level.

            1. user43928 · · focus · HN ↗
              Good point, I now see that they had previously patched the first Artifactory vulnerability, albeit before the HF incident was found.

              Not sure what you mean with the network isolation not being at the VM level. I'm getting a bit lost here in the details of where what network isolation was applied, and the structure of the container runtime and what the apparently compromised parent VM refers to.

              In any case, I understand that after the HF incident was found and they realized the danger, they removed Artifactory and switched to two separate layers of network isolation.

              That said, I believe it was compromised again last week:

              > OpenAI describes an internal research model that, during RL training on September 20, exploited insufficient DNS filtering in its sandbox to contact a public chatbot service. OpenAI’s monitoring flagged the behavior within 15 minutes, a human reviewer acknowledged it three minutes later, and the run was killed about 2.5 hours after the external contact

              Apparently they have now paused training and inference for their most capable model because of this.

              1. charleslmunger · · focus · HN ↗
                My understanding was that they were running on VMs, and there was at least one container inside the VM. So by breaking out of the container they compromised the VM, but they did not break out of the VM itself. The network isolation was enforced at the container level, so they didn't need to break out of the VM and compromise the hypervisor to move laterally. The technical report says:

                >OpenAI also required all Research CaaS workloads to use a micro VM sandbox

                So yeah, they were not using it before and they should have been, and they're doing it they way they should have, now.

                >exploited insufficient DNS filtering in its sandbox

                If your goal is to use dns filtering to prevent your models from reaching the open Internet then you are not really serious about security. People have been tunnelling DNS to get internet access on captive portal wifi since 1998. There's not a lot of details in the blog post but I suspect this was also foreseeable.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.