‹ BackHN Continuity

Thread

Revealing the details of how OpenAI agents hacked Hugging Face

755 points · 472 comments · specked-citrus

  1. GuB-42 · · focus · HN ↗
    So ugly...

    It looks like a primitive chess engine, trying every move, no matter how stupid, until it works. Relying on its ability to do millions of operations rather than having a plan.

    People will try stuff too, but once there is an opening, they will consolidate, generalize, simplify,... before going to the next step. The agents didn't, it is a huge, vaguely directed mess.

    Also, it looked so "loud", querying millions of URL with weird requests. The sandbox as weak as it can get, and there is absolutely zero smart extrusion detection or it would have found it. They used their best AI for attacking, but nothing for protection.

    1. doginasuit · · focus · HN ↗
      This is why I have a very low p(doom). LLMs have an incredible working memory, but they have a hard limit on translating that into good decisions. They get by entirely on their persistence. That works fine in the digital world, but once you cross the boundary into physical space the advantage disappears.
      1. pyronite · · focus · HN ↗
        I don’t know how you quantify a very low p(doom), but this is why mine is high enough to worry me.

        A million AI monkeys at a million AI typewriters, banging away at random, could do amazing damage.

        1. otterley · · focus · HN ↗
          Which will happen first: amazing damage, or reproduce a Shakespeare play?
          1. AnimalMuppet · · focus · HN ↗
            It is easier to destroy than to build.
            1. skinfaxi · · focus · HN ↗
              Is it easier to discover a vulnerability than to introduce one?
              1. AnimalMuppet · · focus · HN ↗
                It is easier to discover existing vulnerabilities and use them to cause massive destruction than it is to plug the existing vulnerabilities.
                1. skinfaxi · · focus · HN ↗
                  You missed my point. Is it easier to discover a novel exploit than it is to build software that is exploitable?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.