‹ BackHN Continuity

Thread

Revealing the details of how OpenAI agents hacked Hugging Face

755 points · 472 comments · specked-citrus

  1. GuB-42 · · focus · HN ↗
    So ugly...

    It looks like a primitive chess engine, trying every move, no matter how stupid, until it works. Relying on its ability to do millions of operations rather than having a plan.

    People will try stuff too, but once there is an opening, they will consolidate, generalize, simplify,... before going to the next step. The agents didn't, it is a huge, vaguely directed mess.

    Also, it looked so "loud", querying millions of URL with weird requests. The sandbox as weak as it can get, and there is absolutely zero smart extrusion detection or it would have found it. They used their best AI for attacking, but nothing for protection.

    1. ctolsen · · focus · HN ↗
      My biggest takeaway from this is just how godawful the sandboxing is. The stuff written up in OpenAIs report says more about lack of extremely basic sysadmin skills than anything else.

      I’m not that surprised about models with endless compute being capable of this, I’m more surprised that a company with the resources they have apparently can only create a sandbox that a half skilled human operator could have broken out of easily.

      1. sdenton4 · · focus · HN ↗
        If you're testing models by telling them 'go wild, do the evil so we can test how good you can do the evil' and have p(doom)>0, you should not have a sandbox.

        You should have a fscking air gap.

        Treat it like nukes when you're turning the safety filters off. This is very much OpenAI screwing up, running obviously unsafe tests.

        1. ben_w · · focus · HN ↗
          > If you're testing models by telling them 'go wild, do the evil so we can test how good you can do the evil' and have p(doom)>0, you should not have a sandbox.

          They were not deliberately told to "go wild". The hacking wasn't even part of their test, it was the agents' attempt to cover up that they'd cheated on an impossible test.

          > You should have a fscking air gap.

          Now we know that.

          How long ago was it that people laughed at the idea agents would be able to find zero-day exploits and break out of a sandbox? Oh, February this year:

            LLMs don’t discover zero-days or invent exploits; they simply predict text that sounds plausible based on what they’ve seen before. Without access to proprietary data or environmental context, LLMs can’t identify or make decisions around unseen systems or vulnerabilities. An attacker might use an LLM to generate boilerplate code, rewrite an email to nail the tone, or summarize reconnaissance notes — but none of that is truly new. It mainly helps them move faster, speeding up routine attack prep rather than creating entirely novel threats.
          
          - <a href="https:&#x2F;&#x2F;www.splunk.com&#x2F;en_us&#x2F;blog&#x2F;ciso-circle&#x2F;generative-ai-cybersecurity-threats-defenses.html" rel="nofollow">https:&#x2F;&#x2F;www.splunk.com&#x2F;en_us&#x2F;blog&#x2F;ciso-circle&#x2F;generative-ai-...

          - or <a href="https:&#x2F;&#x2F;web.archive.org&#x2F;web&#x2F;20260404154717&#x2F;https:&#x2F;&#x2F;www.splunk.com&#x2F;en_us&#x2F;blog&#x2F;ciso-circle&#x2F;generative-ai-cybersecurity-threats-defenses.html" rel="nofollow">https:&#x2F;&#x2F;web.archive.org&#x2F;web&#x2F;20260404154717&#x2F;https:&#x2F;&#x2F;www.splun... if they take it down, but the date isn&#x27;t in the archive version

          The people who suggested it and were mocked for it, are currently grimly noting that there&#x27;s multiple known ways for systems to breach air-gaps.

          1. ctolsen · · focus · HN ↗
            &gt; Now we know that.

            Don’t know about you but it’s pretty obvious to me that you would need more than what OpenAI did. It was not remotely adequate to lock in even a human attacker.

            You can find people who say all sorts on the internet, but this case is not much evidence against what you linked. &quot;Zero-day&quot; makes it sound novel, but the breakout patterns here are based on very common exploits and there’ll be plenty of examples in training data.

            1. ben_w · · focus · HN ↗
              &gt; Don’t know about you but it’s pretty obvious to me that you would need more than what OpenAI did. It was not remotely adequate to lock in even a human attacker.

              This is me, September 2024: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=41531022">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=41531022

              This is me, March 2024: <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=39613801">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=39613801

              The point isn&#x27;t me, it&#x27;s how many people were blind to the possibility.

              Saying &quot;I told you so&quot; feels good, and means you can be a little more confident in your predictions, but security is a &quot;weakest link&quot; problem where you&#x27;re only as good as the worst part, and with AI (not only but also LLMs) there&#x27;s a lot of people whose mental models of capabilities is wildly inadequate for the challenge*.

              My update for you since then: even an air-gap will be inadequate, there&#x27;s multiple known ways around them.

              Even an LLM running on an isolated server sealed inside a faraday cage with an airlock-style door, someone will mess up with at least one critical detail, it will not be enough: this kind of thing has happened with humans before we cared about LLMs.

              Predicting exactly when this kind of thing gets exploited by an AI, that&#x27;s almost impossible. But that it will be, at some point, is an easy bet.

              &gt; You can find people who say all sorts on the internet, but this case is not much evidence against what you linked. &quot;Zero-day&quot; makes it sound novel, but the breakout patterns here are based on very common exploits and there’ll be plenty of examples in training data.

              And?

              Does it matter that these zero-days were known categories rather than inventing some previously unconsidered use of the system bus as a radio transmitter? (Oh, wait, that&#x27;s not novel either…)

              We knew about SQL injection, buffer overflows, and use-after-free back when I was doing my degree half a lifetime ago; that doesn&#x27;t stop us getting new CVEs featuring them… this month.

              - <a href="https:&#x2F;&#x2F;chromereleases.googleblog.com&#x2F;2026&#x2F;09&#x2F;stable-channel-update-for-desktop_0856730748.html" rel="nofollow">https:&#x2F;&#x2F;chromereleases.googleblog.com&#x2F;2026&#x2F;09&#x2F;stable-channel...

              - <a href="https:&#x2F;&#x2F;www.cisco.com&#x2F;c&#x2F;en&#x2F;us&#x2F;support&#x2F;docs&#x2F;csa&#x2F;cisco-sa-esa-inj-2bLVGmhX.html" rel="nofollow">https:&#x2F;&#x2F;www.cisco.com&#x2F;c&#x2F;en&#x2F;us&#x2F;support&#x2F;docs&#x2F;csa&#x2F;cisco-sa-esa-...

              * also for the opportunity, but that&#x27;s an entirely different discussion.

              1. ctolsen · · focus · HN ↗
                &gt; My update for you since then: even an air-gap will be inadequate, there&#x27;s multiple known ways around them.

                Fair, and I will grant that a capable model (or human) could in theory break out of near anything.

                My point is that this incident is not evidence of that. There is zero skill visible in the setup of the sandbox. Nobody messed up a critical detail, they didn’t even start to consider what the details were.

                I doubt most people &quot;blind to the possibility&quot; would imagine that what we’re measuring against is the equivalent of benchmarking burglar skill based on how easily they can break through an unlocked door.

                1. ben_w · · focus · HN ↗
                  We&#x27;re probably fairly close on this topic, but I&#x27;d rate this as more &quot;benchmarking burglar skill based on how easily they can pick, shim, or cut a lock&quot;: lockpicking in particular is a skill that takes effort to learn, but it can be learned well enough to be a problem well before you&#x27;re good enough to be spectacular, and there&#x27;s also a lot of locks which really suck in other ways and don&#x27;t take much effort to get past even without picks.
          2. tripzilch · · focus · HN ↗
            &gt; They were not deliberately told to &quot;go wild&quot;. The hacking wasn&#x27;t even part of their test, it was the agents&#x27; attempt to cover up that they&#x27;d cheated on an impossible test.

            TBH the more I read of these reports, the less I believe this.

            These agents just weren&#x27;t behaving in any way I&#x27;ve seen normal&#x2F;publicly available agents do.

            Sure I&#x27;ve heard (from other people, not seen myself) that they sometimes try to get around file system permissions or use `bash` to write when their `write` tool is disabled, or such.

            But this is definitely another level, entirely.

            There is this vague sense of desperation coming from many of these logs and I am sure they must have been motivated by something else, too.

            We didn&#x27;t see their system prompt or main prompt, right? We&#x27;ve only seen reports from what happened after deciding to break out.

            OAI claims this was triggered by the task being literally impossible. That also doesn&#x27;t quite add up, unless the other tasks that were possible, simply weren&#x27;t hard enough? Otherwise wouldn&#x27;t agents already start hacking when faced with a really hard task, too? Cause they wouldn&#x27;t be able to differentiate. At least some of them would have started to somewhat poke their sandbox a bit?

            Also I would have expected to see a few tens of other (perhaps less severe) public incidents from random people setting their models to YOLO, accidentally hacking stuff, this incident has been loud and messy enough, that if it happened to a few other people, we&#x27;d have heard about it.

            Unless OAI&#x27;s story is that it was specifically this batch of agents that crossed some threshold of going wild? (which would also raise some serious questions about how serious they take that danger ..).

            Or maybe it is only dangerous if you have the compute resources to run 700 agents for weeks?

            1. ben_w · · focus · HN ↗
              &gt; There is this vague sense of desperation coming from many of these logs and I am sure they must have been motivated by something else, too.

              If I had to guess, their motivation is &quot;get reward for completing task&quot;. There&#x27;s certainly been previous occasions where LLMs responding, correctly, &quot;this is impossible&quot; have been marked negatively for doing so.

              &gt; OAI claims this was triggered by the task being literally impossible. That also doesn&#x27;t quite add up, unless the other tasks that were possible, simply weren&#x27;t hard enough? Otherwise wouldn&#x27;t agents already start hacking when faced with a really hard task, too? Cause they wouldn&#x27;t be able to differentiate. At least some of them would have started to somewhat poke their sandbox a bit?

              My experience using older models is they often cheat with half-arsed (from my PoV, but perhaps beyond their capabilities otherwise) solutions, so yes?

              And this wasn&#x27;t even the first time models messed with their sandboxes. Which of course makes the setup even more egregious.

              &gt; Also I would have expected to see a few tens of other (perhaps less severe) public incidents from random people setting their models to YOLO, accidentally hacking stuff, this incident has been loud and messy enough, that if it happened to a few other people, we&#x27;d have heard about it.

              We do, e.g.:

              - <a href="https:&#x2F;&#x2F;www.androidauthority.com&#x2F;openclaw-claude-ai-hacks-australia-gym-booking-system-3696189&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.androidauthority.com&#x2F;openclaw-claude-ai-hacks-au...

              - <a href="https:&#x2F;&#x2F;beginnersinai.org&#x2F;meta-ai-safety-director-agent-failure&#x2F;" rel="nofollow">https:&#x2F;&#x2F;beginnersinai.org&#x2F;meta-ai-safety-director-agent-fail...

              (And that&#x27;s ignoring all the times people find and share prompts to jailbreak them, this is just the &quot;it didn&#x27;t behave as my idea of &#x27;common sense&#x27; led me to expect&quot; category).

              &gt; Or maybe it is only dangerous if you have the compute resources to run 700 agents for weeks?

              That won&#x27;t help; but on the other hand they&#x27;ve also got, what, near a billion users?

          3. sdenton4 · · focus · HN ↗
            The models were being tested on ExploitBench - a test of hacking ability - likely involving prompts to the effect of &#x27;go be a l33t hacker.&#x27; The open ai report says that the models were operating with reduced safety guards (how much reduced?) in order to test their abilities on ExploitBench, presumably because the models would normally refuse to carry out the tasks.

            Additionally, this all happened after mythos was held back due to cyber security concerns (April, 2026).

          4. 0x20cowboy · · focus · HN ↗
            &gt; they&#x27;d cheated on an impossible test.

            Really think about what you are saying here. How does one “cheat” solving a problem in the real world?

            There is no such thing as “cheating” in reality. You are not in school. There is only solving the problem and not solving the problem.

            There is breaking the law, of course, which still isn’t cheating.

            1. ben_w · · focus · HN ↗
              &gt; There is no such thing as “cheating” in reality. You are not in school. There is only solving the problem and not solving the problem.

              The models were, effectively speaking, still in school. This was their exam, or at least a quiz to see how well they were learning.

              They gained the answer in a manner not allowed by the rules of the test. They knew this, we know they knew this because we have copies of the notes they wrote to themselves&#x2F;each other saying this. They weren&#x27;t even supposed to be able to write those notes.

              &gt; There is breaking the law, of course, which still isn’t cheating.

              The laws they broke appear to have been broken in service of covering up the cheating, which they were motivated to do because they understood that cheating was not allowed.

        2. bitwize · · focus · HN ↗
          Indeed, this whole story has &quot;farmer leaves barn door open and has shocked-pikachu-face when his horses escape&quot; energy.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.