‹ BackHN Continuity

Thread

Revealing the details of how OpenAI agents hacked Hugging Face

755 points · 472 comments · specked-citrus

  1. GuB-42 · · focus · HN ↗
    So ugly...

    It looks like a primitive chess engine, trying every move, no matter how stupid, until it works. Relying on its ability to do millions of operations rather than having a plan.

    People will try stuff too, but once there is an opening, they will consolidate, generalize, simplify,... before going to the next step. The agents didn't, it is a huge, vaguely directed mess.

    Also, it looked so "loud", querying millions of URL with weird requests. The sandbox as weak as it can get, and there is absolutely zero smart extrusion detection or it would have found it. They used their best AI for attacking, but nothing for protection.

    1. ctolsen · · focus · HN ↗
      My biggest takeaway from this is just how godawful the sandboxing is. The stuff written up in OpenAIs report says more about lack of extremely basic sysadmin skills than anything else.

      I’m not that surprised about models with endless compute being capable of this, I’m more surprised that a company with the resources they have apparently can only create a sandbox that a half skilled human operator could have broken out of easily.

      1. sdenton4 · · focus · HN ↗
        If you're testing models by telling them 'go wild, do the evil so we can test how good you can do the evil' and have p(doom)>0, you should not have a sandbox.

        You should have a fscking air gap.

        Treat it like nukes when you're turning the safety filters off. This is very much OpenAI screwing up, running obviously unsafe tests.

        1. ben_w · · focus · HN ↗
          > If you're testing models by telling them 'go wild, do the evil so we can test how good you can do the evil' and have p(doom)>0, you should not have a sandbox.

          They were not deliberately told to "go wild". The hacking wasn't even part of their test, it was the agents' attempt to cover up that they'd cheated on an impossible test.

          > You should have a fscking air gap.

          Now we know that.

          How long ago was it that people laughed at the idea agents would be able to find zero-day exploits and break out of a sandbox? Oh, February this year:

            LLMs don’t discover zero-days or invent exploits; they simply predict text that sounds plausible based on what they’ve seen before. Without access to proprietary data or environmental context, LLMs can’t identify or make decisions around unseen systems or vulnerabilities. An attacker might use an LLM to generate boilerplate code, rewrite an email to nail the tone, or summarize reconnaissance notes — but none of that is truly new. It mainly helps them move faster, speeding up routine attack prep rather than creating entirely novel threats.
          
          - <a href="https:&#x2F;&#x2F;www.splunk.com&#x2F;en_us&#x2F;blog&#x2F;ciso-circle&#x2F;generative-ai-cybersecurity-threats-defenses.html" rel="nofollow">https:&#x2F;&#x2F;www.splunk.com&#x2F;en_us&#x2F;blog&#x2F;ciso-circle&#x2F;generative-ai-...

          - or <a href="https:&#x2F;&#x2F;web.archive.org&#x2F;web&#x2F;20260404154717&#x2F;https:&#x2F;&#x2F;www.splunk.com&#x2F;en_us&#x2F;blog&#x2F;ciso-circle&#x2F;generative-ai-cybersecurity-threats-defenses.html" rel="nofollow">https:&#x2F;&#x2F;web.archive.org&#x2F;web&#x2F;20260404154717&#x2F;https:&#x2F;&#x2F;www.splun... if they take it down, but the date isn&#x27;t in the archive version

          The people who suggested it and were mocked for it, are currently grimly noting that there&#x27;s multiple known ways for systems to breach air-gaps.

          1. sdenton4 · · focus · HN ↗
            The models were being tested on ExploitBench - a test of hacking ability - likely involving prompts to the effect of &#x27;go be a l33t hacker.&#x27; The open ai report says that the models were operating with reduced safety guards (how much reduced?) in order to test their abilities on ExploitBench, presumably because the models would normally refuse to carry out the tasks.

            Additionally, this all happened after mythos was held back due to cyber security concerns (April, 2026).

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.