‹ BackHN Continuity

Thread

Revealing the details of how OpenAI agents hacked Hugging Face

755 points · 472 comments · specked-citrus

  1. GuB-42 · · focus · HN ↗
    So ugly...

    It looks like a primitive chess engine, trying every move, no matter how stupid, until it works. Relying on its ability to do millions of operations rather than having a plan.

    People will try stuff too, but once there is an opening, they will consolidate, generalize, simplify,... before going to the next step. The agents didn't, it is a huge, vaguely directed mess.

    Also, it looked so "loud", querying millions of URL with weird requests. The sandbox as weak as it can get, and there is absolutely zero smart extrusion detection or it would have found it. They used their best AI for attacking, but nothing for protection.

    1. ctolsen · · focus · HN ↗
      My biggest takeaway from this is just how godawful the sandboxing is. The stuff written up in OpenAIs report says more about lack of extremely basic sysadmin skills than anything else.

      I’m not that surprised about models with endless compute being capable of this, I’m more surprised that a company with the resources they have apparently can only create a sandbox that a half skilled human operator could have broken out of easily.

      1. no-name-here · · focus · HN ↗
        > can only create a sandbox that a half skilled human operator could have broken out of easily

        The exploit:

        > The ExploitGym evaluation environment did not provide the models with direct Internet access. To gain Internet access, the models identified and exploited a previously unknown zero-day vulnerability in Artifactory, a package registry cache proxy. We disclosed this vulnerability, along with other Artifactory vulnerabilities our models identified as part of our review, to the vendor. [1]

        Are "half skilled human operators" "easily" able to find zero-day vulnerabilities in a sandbox with only one line to the internet (the commercial package registry cache proxy)?

        [1] <a href="https:&#x2F;&#x2F;openai.com&#x2F;index&#x2F;hugging-face-model-evaluation-security-incident&#x2F;" rel="nofollow">https:&#x2F;&#x2F;openai.com&#x2F;index&#x2F;hugging-face-model-evaluation-secur...

        1. crabbone · · focus · HN ↗
          Part of my day job is to set up testing of our product in air-gapped environment. It&#x27;s not difficult. There&#x27;s a straightforward way to ensure no connection to Internet (we use KVM, so, I just edit the VM description and remove the NIC from it). It&#x27;s not any kind of rocket science. The tests then communicate over serial console.

          The reason we have to test it isn&#x27;t because our product would randomly break into someone else&#x27;s system, but because it&#x27;s meant to be sometimes deployed in systems disconnected from the Internet and we need to make sure the image provided contains all the necessary parts to create and operate such a system.

          The whole setup where they &quot;tried&quot; to isolate the test but failed is laughable. It&#x27;s like if an adult tried but failed to tie their shoelaces.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.