‹ BackHN Continuity

Thread

Revealing the details of how OpenAI agents hacked Hugging Face

755 points · 472 comments · specked-citrus

  1. GuB-42 · · focus · HN ↗
    So ugly...

    It looks like a primitive chess engine, trying every move, no matter how stupid, until it works. Relying on its ability to do millions of operations rather than having a plan.

    People will try stuff too, but once there is an opening, they will consolidate, generalize, simplify,... before going to the next step. The agents didn't, it is a huge, vaguely directed mess.

    Also, it looked so "loud", querying millions of URL with weird requests. The sandbox as weak as it can get, and there is absolutely zero smart extrusion detection or it would have found it. They used their best AI for attacking, but nothing for protection.

    1. doginasuit · · focus · HN ↗
      This is why I have a very low p(doom). LLMs have an incredible working memory, but they have a hard limit on translating that into good decisions. They get by entirely on their persistence. That works fine in the digital world, but once you cross the boundary into physical space the advantage disappears.
      1. pyronite · · focus · HN ↗
        I don’t know how you quantify a very low p(doom), but this is why mine is high enough to worry me.

        A million AI monkeys at a million AI typewriters, banging away at random, could do amazing damage.

        1. tharkun__ · · focus · HN ↗
          Especially when they cross into the physical realm as in not properly secured and air gapped control systems. SCADA is scary.
          1. mrob · · focus · HN ↗
            That lowers P(doom), because it gives AI a chance to do enough damage to make people take the threat seriously before anybody gets recursive self-improvement working.
            1. doginasuit · · focus · HN ↗
              Exactly, there's no path to AI reaching that level of dominance without taking actions with high stakes.
            2. asdff · · focus · HN ↗
              The thing is we are basically guaranteeing this to happen. We might kill off all the models that seem like they are going to threaten the power structure of the planet through these sorts of things. That will work for a while. But just like most things in life, by sheer dumb random chance, there will be once case that manages to have some way to evade detection, proliferate, then dominate. We are basically giving it selective pressure to favor this outcome.
          2. phil21 · · focus · HN ↗
            > Especially when they cross into the physical realm as in not properly secured and air gapped control systems. SCADA is scary.

            What about the bad actors (choose your own evildoer here) who purposefully do not air gap their agents? And specifically train them to attack in such a manner?

            I'd much rather have relatively benign stuff like this hit first, because the former is coming sooner than later. It's already here in a limited manner, likely more than any of us currently realize.

            Botnets could crack passwords faster than anyone thought possible over 20 years ago now. This is just the latest iteration of such a concept.

            There is so much low hanging fruit in this space that frontier models are currently utterly irrelevant. It's going to take decades of human-speed securing of IT to make superintelligence or whatever you want to call it a necessary component for such attacks.

            At this point, someone with a rack or three of GPUs with 100kw to burn can replicate such attacks if they feel like it. the bar for entry is not even 7 figures.

          3. antii · · focus · HN ↗

            [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.