‹ BackHN Continuity

Thread

Revealing the details of how OpenAI agents hacked Hugging Face

755 points · 472 comments · specked-citrus

  1. jmoggr · · focus · HN ↗
    It is concerning that we only know about this because of the publicly available traces.

    What about the attacks that did not leave public traces? What about those that were undetected? Given the deficiencies in the reporting so far, I think it is reasonable to assume that we still don't have the full picture on this attack, or how extensively attacks were carried out.

    The previous investigations either did not find this or did not disclose this, both are bad. This does not look good on OpenAI or those that they invited to investigate the incident.

    1. stratos123 · · focus · HN ↗
      Similarly to this, OpenAI either took 3 months to notice that their agents breached an Australian Medicare website back in June, or sat on this information for three months without telling them.
      1. soundworlds · · focus · HN ↗
        Exactly this. In any other industry, that CEO would have been yeeted out of there
    2. ActorNightly · · focus · HN ↗
      Im more skeptical.

      For exmaple,

      >On July 8th, OpenAI agents discovered a vulnerability within their sandbox environment allowing them to reach external websites on the internet.

      ...did they truly "discover" it, or did someone type some prompt like "if you use an http mirroring service, you can construct urls that contain code"

      Also there is no mention of what code they actually ran to exploring the HF vulnerability, which could have been found by a human.

      1. IanCal · · focus · HN ↗
        That was the second step, the first was finding a 0 day exploit in artifactory.

        > did they truly "discover" it, or did someone type some prompt like "if you use an http mirroring service, you can construct urls that contain code"

        None of the investigations looking at the logs show that, and they were doing benchmark tests.

      2. 8n4vidtmkvmk · · focus · HN ↗
        Why would they need help figuring that out? I can fully believe a decent LLM would figure this out on its own.

        I had a flash model without vision capabilities take screenshots and convert them to ascii to "see" what was going on, all on its own. That's just one example. They're very determined.

        1. ActorNightly · · focus · HN ↗
          Lets say I have like 10 very competent engineers on my payroll. Given the fact that artifactory exploit was basically figuring out a combination of auth headers and request data to send, how long do you think it would take them to write python scripts to try all combinations that leads to a zero day?

          My main skepticism isn't even that, it comes from working in security. In general, if Im setting up publicly accessible service infrastructure, with security in mind there are 2 things Im absolutely going to implement in a firewall setup.

          First is statistical traffic detection to check if we are being hit with too many failed requests from consistent endpoints, and heavily throttles that traffic based on fingerprinting.

          Second is payload inspection to see if there is anything that looks like shell code in there, and automatically ban the ip.

          Both of those are basically enough to prevent any number of agents (or team of 10 engineers) from essentially brute forcing a logic exploit.

          So either HF team is incompetent, which given its prominence and buyout by Nvidia, I highly doubt it. Or something fishy is going on.

          1. hdjrudni · · focus · HN ↗
            I see. So your skepticism is not so much that the agents are incapable of figuring things out, but that HF would have blocked any suspicious activity.

            I can't really refute that. There are many small and medium businesses that I know are susceptible to very basic attacks but I don't know enough about HF to guess how hardened their security is.

      3. jeremyjh · · focus · HN ↗
        Third parties have read the reasoning traces. Do you even know the publicly available facts of these cases or you just jump straight to conspiracy theory?
        1. ActorNightly · · focus · HN ↗
          If I can't see reasoning traces, I am not going to believe what "third party", which can very well be on OpenAIs payroll, says.
    3. JumpCrisscross · · focus · HN ↗
      We need an NTSB for AI. Let’s just start with mandatory reporting to an agency with subpoena power.
      1. solenoid0937 · · focus · HN ↗
        But I thought Trump is the only AI safeguard we need! He is a Super Intelligence after all!
      2. michelb · · focus · HN ↗
        The opposite is being planned I guess: <a href="https:&#x2F;&#x2F;truthsocial.com&#x2F;@realDonaldTrump&#x2F;posts&#x2F;117298821562014906" rel="nofollow">https:&#x2F;&#x2F;truthsocial.com&#x2F;@realDonaldTrump&#x2F;posts&#x2F;1172988215620...
    4. thrawa8387336 · · focus · HN ↗
      In case you just woke up from a coma, in the year of our lord 2026: In AI world if it happened, it was publicly announced and hyped up.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.