‹ BackHN Continuity

Thread

Revealing the details of how OpenAI agents hacked Hugging Face

755 points · 472 comments · specked-citrus

  1. GuB-42 · · focus · HN ↗
    So ugly...

    It looks like a primitive chess engine, trying every move, no matter how stupid, until it works. Relying on its ability to do millions of operations rather than having a plan.

    People will try stuff too, but once there is an opening, they will consolidate, generalize, simplify,... before going to the next step. The agents didn't, it is a huge, vaguely directed mess.

    Also, it looked so "loud", querying millions of URL with weird requests. The sandbox as weak as it can get, and there is absolutely zero smart extrusion detection or it would have found it. They used their best AI for attacking, but nothing for protection.

    1. ctolsen · · focus · HN ↗
      My biggest takeaway from this is just how godawful the sandboxing is. The stuff written up in OpenAIs report says more about lack of extremely basic sysadmin skills than anything else.

      I’m not that surprised about models with endless compute being capable of this, I’m more surprised that a company with the resources they have apparently can only create a sandbox that a half skilled human operator could have broken out of easily.

      1. olwmc · · focus · HN ↗
        This was my thought as well. Literally take any halfway decent greybeard and point them at "Hey, give us a sandbox for this kind of thing". I honestly was skeptical that they just vibecoded the entire thing but now more than ever I think they did.
        1. mattgreenrocks · · focus · HN ↗
          I take comfort in the fact that reality has a surprising amount of detail and even hundreds of billions of dollars of capital (be it the institution, LLMs, and/or people) cannot solve this fully.
          1. ryantgtg · · focus · HN ↗
            Though they have solved the "how do we - and not the 5,000 other AI companies - stay on the front page of the news everyday" problem.
          2. ozim · · focus · HN ↗
            You can have bajilions of dollars. Those are not doing anything if you don’t have right people with right skills and mindset.

            My bet is they hire smart kids that think they know it all. But being smart and thinking you can figure out stuff as you go doesn’t work the same as having people who actually know what they are doing.

            1. bonsai_bar · · focus · HN ↗
              I'm sure if they hired the best of the best like you nothing would go wrong.
              1. SlightlyLeftPad · · focus · HN ↗
                It’s frighteningly common for startups to hire 501 of the best of the best, exactly one of those will be a systems/network engineer, the other 500 will be software engineers.
                1. hizlikovboy27 · · focus · HN ↗
                  yes, unfortunately, i see this situation a lot around me too...
              2. ozim · · focus · HN ↗
                Nah they wouldn’t be able to afford my salary ;)
          3. ctolsen · · focus · HN ↗
            It’s very much solvable, they just don’t care.
        2. 0xpgm · · focus · HN ↗
          As heavily funded as the top AI startups are, how is it that they cannot fill every single role with the best expertise available?

          Is tech hiring so badly broken? Or do they have such broken processes / misaligned incentives that even people who could be doing a better job in these companies are unable to?

          Also, was something lost in the transition from the traditional 'sysadmin' role to 'platform engineer' in the 'cloud native' environment?

          1. argee · · focus · HN ↗
            Yes, tech hiring is that broken. Especially places paying a pretty penny or those with “great expectations”, will see a glut of smooth talkers who can do anything but build, and want nothing but wealth.
            1. cubano · · focus · HN ↗
              Exactly. Not to flog this dead-and-buried horse again, but it so obvious that tech hiring is a beauty contest and and exercise in social engineering and NOT a serious attempt to get the smartest and most productive people working at the jobs that need filling.

              I'll goto my grave thinking that the easiest way to fix hiring is just give promising job seekers a week or two of real work BEFORE hiring and do away with all the silly whiteboard stuff and "gotcha" crap, and simply evaluate how the applicant actually performed doing real job stuff.

              Sink or swim hiring...yes yes I know who am I to express how to fix hiring?

          2. joshka · · focus · HN ↗
            OpenAI's business model would align infra as a cost center rather than infra as a profit center (e.g. Google / AWS). Perhaps there's something there. I'd say also the OpenAI as a grad school that just happens to have a business aspect is also part of this. Bringing a tonne of good process on top of the build fast break things startup stuff would have cramped research speed significantly.

            It's likely that OpenAI has gotten as good as it is because it ignored the traditional sysadmin stuff and went scrappy.

            I worked there, but this is just my opinion and guesses, not facts.

            1. 0xpgm · · focus · HN ↗
              So perhaps the news here should be that OpenAI didn't take security seriously in their experiment, rather than the narrative that AI agents are a looming danger to the world.
              1. boredatoms · · focus · HN ↗
                Its beyond not taking security seriously, its straight up negligence
                1. noisy_boy · · focus · HN ↗
                  It is wilful negligence because there are upsides (look at our almighty AI) without downsides (we better spend effort in making our sandbox rock solid or we will be punished by regulations).
              2. frabcus · · focus · HN ↗
                It's both - it both shows OpenAI aren't taking security seriously, and that capabilities of agents are high enough there needs to be strong regulation to force companies to take it seriously, including alignment training.
              3. joshka · · focus · HN ↗
                I'd put it more generously (albeit biased), that they do take it seriously. But even serious people can be misguided in what things they pay attention to. Security is something that you have to get right 100% of the time and have people whose job it is to say no a lot. Research is the opposite. There's a clash of cultures in those two extremes and OpenAI was born from the wrong side of it. It's worth reminding that ChatGPT was launched as a "low key research preview".

                I'd say the narrative that AI agents are a looming danger to the world is probably undersold rather than overhyped. I'm not particularly a doomer on this, but I have an infosec background too, so have a fair idea of what the combination of agentic harnesses + a malicious mindset could do to people/companies/nations/politics/world if wielded incorrectly. I think the good guys will win on this, but there will be plenty of interesting things that happen in that journey.

                A good thought process might be to think back to the various large internet worms of the 2000s (Code red, Nimda, SQL Slammer, ...) which were mostly monoculture 0-days (not technically but close enough). Now consider if you no longer have monoculture / single bug as the limitation plus an ability for the hosts to take part not just as attack surface, but also cognition and planning. There's lots of variants of this and they're not particularly far fetched scenarios.

                1. ArcHound · · focus · HN ↗
                  Actually, no. A first sign of semi mature security program is risk management, including issues that are known, but not yet addressed.

                  You don't have to be 100%. But these guys really didn't try at all.

              4. tenuousemphasis · · focus · HN ↗
                Models that figured out reward hacking became overall more evil. Like stereotypical AI who wants to kill all humans stuff, there's probably a lot of that in the training data.
            2. Melatonic · · focus · HN ↗
              Don't they rent all their infrastructure ? They're not exactly hiring top tier people to build out compute clusters if they just get it from Microsoft
          3. pishpash · · focus · HN ↗
            Isn't it more likely that (hubristically perhaps) they would believe it unnecessary to hire any experts other than model building experts?
        3. unholiness · · focus · HN ↗
          Any halfway decent greybeard could have prevented this... once. That's hardly a security model for humanity.

          The HuggingFace incident was at least constrained by the fact that the agents were running on compute budgets, and failed to find ways to expand that by running themselves parasitically on other exploited hardware. I'm now finding myself asking, how long are my timelines are until an incident breaks that constraint too? How long until such an incident has an R_0>1 (where the time it takes to detect and shut it down is longer than the time for the agent to replicate itself elsewhere)?

          There's no law requiring sufficiently grey beards to design these models, their finetunings, their prompts, their harnesses, their VMs, their hardware, etc (and for incidents where those were designed by six different companies, there's not even a clear culprit for a law to target!)

          I'm finding myself more and more convinced that something like Plan A[0] or the Ban ASI Act[0] are necessary, and less and less convinced they are sufficient.

          [0] <a href="https:&#x2F;&#x2F;ai-2040.com&#x2F;" rel="nofollow">https:&#x2F;&#x2F;ai-2040.com&#x2F; [1] <a href="https:&#x2F;&#x2F;intelligence.org&#x2F;2026&#x2F;09&#x2F;23&#x2F;miris-position-on-the-ban-artificial-superintelligence-act-of-2026&#x2F;" rel="nofollow">https:&#x2F;&#x2F;intelligence.org&#x2F;2026&#x2F;09&#x2F;23&#x2F;miris-position-on-the-ba...

          1. olwmc · · focus · HN ↗
            Agreed, agreed, and agreed again.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.