‹ BackHN Continuity

Thread

Revealing the details of how OpenAI agents hacked Hugging Face

755 points · 472 comments · specked-citrus

  1. GuB-42 · · focus · HN ↗
    So ugly...

    It looks like a primitive chess engine, trying every move, no matter how stupid, until it works. Relying on its ability to do millions of operations rather than having a plan.

    People will try stuff too, but once there is an opening, they will consolidate, generalize, simplify,... before going to the next step. The agents didn't, it is a huge, vaguely directed mess.

    Also, it looked so "loud", querying millions of URL with weird requests. The sandbox as weak as it can get, and there is absolutely zero smart extrusion detection or it would have found it. They used their best AI for attacking, but nothing for protection.

    1. jbrooks84 · · focus · HN ↗
      Yup literally no security and they wonder how they got out
      1. no-name-here · · focus · HN ↗
        > literally no security

        What is the source that there was "literally no security"?

        > and they wonder how they got out

        OpenAI publicly announced months ago how the model got out:

        > The ExploitGym evaluation environment did not provide the models with direct Internet access. To gain Internet access, the models identified and exploited a previously unknown zero-day vulnerability in Artifactory, a package registry cache proxy. We disclosed this vulnerability, along with other Artifactory vulnerabilities our models identified as part of our review, to the vendor. [1]

        [1] <a href="https:&#x2F;&#x2F;openai.com&#x2F;index&#x2F;hugging-face-model-evaluation-security-incident&#x2F;" rel="nofollow">https:&#x2F;&#x2F;openai.com&#x2F;index&#x2F;hugging-face-model-evaluation-secur...

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.