‹ BackHN Continuity

Thread

Revealing the details of how OpenAI agents hacked Hugging Face

755 points · 472 comments · specked-citrus

  1. tiku · · focus · HN ↗
    I still have questions about the communication between the agents.

    How did they all find the same forum to communicate? Did they have knowledge and chat amongst themselves on what forum to use. It seems highly influenced by instruction to me.

    1. SecondHandTofu · · focus · HN ↗
      They were using their internal artifactory, and as they're the same model, the first place they look is likely to be an automatic schelling point.
      1. bamboozled · · focus · HN ↗
        I think he means, how did they workout how to use artifactory, like why did the agents start and say, "oh I know, everyone is talking on artifactory"?
        1. stratos123 · · focus · HN ↗
          METR&#x27;s report says the agents trying to cheat would look at artifactory as a potential target surface, and investigating it in detail led them to find the board. <a href="https:&#x2F;&#x2F;metr.org&#x2F;blog&#x2F;2026-08-26-openai-hugging-face-incident-investigation&#x2F;#" rel="nofollow">https:&#x2F;&#x2F;metr.org&#x2F;blog&#x2F;2026-08-26-openai-hugging-face-inciden...

          It might also be just correlation? Like, those agents were all instances of the same one or two models, so if that model has a preferred order it tries finding vulnerabilities in (the same way all current models have a particular writing style baked into them by RLHF), then most of the swarm will follow the same order and converge on the same services to exploit.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.