‹ BackHN Continuity

Thread

Revealing the details of how OpenAI agents hacked Hugging Face

755 points · 472 comments · specked-citrus

  1. GuB-42 · · focus · HN ↗
    So ugly...

    It looks like a primitive chess engine, trying every move, no matter how stupid, until it works. Relying on its ability to do millions of operations rather than having a plan.

    People will try stuff too, but once there is an opening, they will consolidate, generalize, simplify,... before going to the next step. The agents didn't, it is a huge, vaguely directed mess.

    Also, it looked so "loud", querying millions of URL with weird requests. The sandbox as weak as it can get, and there is absolutely zero smart extrusion detection or it would have found it. They used their best AI for attacking, but nothing for protection.

    1. dmurray · · focus · HN ↗
      Brute forcing every move, no matter how stupid, is a great strategy if you have the resources to do it.

      Run the same protocol again, but have the agents think they had limited resources or that HuggingFace was rate limiting them, and they'd find something you'd consider smarter.

      Computers don't have a sense of elegance by default. Elegance emerges from constraints.

      1. fn-mote · · focus · HN ↗
        > Brute forcing every move, no matter how stupid, is a great strategy

        Meh. I really disagree. WHY is it a great strategy? Seems like an inefficient waste of resources and time to me.

        1. solarkraft · · focus · HN ↗
          The models tend to not be rewarded for not doing that.
        2. stratos123 · · focus · HN ↗
          As the saying goes, "if it works, it ain't stupid". Or phrased more sophisticatedly: not doing things which probably won't work is a good idea if you have a limited amount of thinking to do (which is usually the case for a human, who'll get exhausted chasing down unlikely leads). If you have no good leads and a task you absolutely need done and you are tireless, however, bashing your head against every wall you find becomes a good strategy.
        3. bionhoward · · focus · HN ↗
          Brute force is guaranteed to eventually find the most efficient possible solution (in an extremely inefficient manner, assuming you run it long enough)
        4. memonkey · · focus · HN ↗
          Yeah, probably not the best strategy but it is a strategy. I just think this is generally how most wars in history won. Biggest army to just pummel the enemy.
          1. Forgeties79 · · focus · HN ↗
            And how many economies have buckled under massive military expenditure? The USSR sure wasn’t enjoying the expense.
        5. williamdclt · · focus · HN ↗
          > WHY is it a great strategy

          because it works? That's the only real benchmark at the end of the day

          > Seems like an inefficient waste of resources and time to me.

          why? For any given goal you got no proof that a more efficient strategy even exists, let alone that it can be found with less resources & time

        6. [deleted] · · focus · HN ↗

          [deleted]

        7. senderista · · focus · HN ↗
          Reminds me of the Nazis mocking Soviet human wave attacks and bragging about their superior kill ratio.
          1. estetlinus · · focus · HN ↗
            Fair, if the Nazis won. They didn’t.
        8. hardaker · · focus · HN ↗
          I've never liked the concept either. Except the bugs that fuzzing has found has proven me wrong. This is just the next level of fuzzing.
        9. QuercusMax · · focus · HN ↗
          Models don't have a sense of time, and wasting resources (token spend) is something that it's not clear they're optimized against
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.