‹ BackHN Continuity

Thread

OpenAI agents tried to bruteforce a UN website's API fields

85 points · 87 comments · intunderflow

  1. cmiles8 · · focus · HN ↗
    The more of these that come out the more incompetent OpenAI looks. It would appear there was a total lack of basic controls in place for running these tests.
    1. Legend2440 · · focus · HN ↗
      I think they did not expect that models were capable of this level of sandbox escape (prior models certainly didn't have this kind of agency) and weren't prepared.

      All these incidents happened between April-July 2026; prior to that, models weren't capable yet. And after that, they were aware and watching much more closely.

      1. schainks · · focus · HN ↗
        > they were aware and watching much more closely.

        I've love to know the reason they never considered air gapping systems before the models got powerful enough.

        It's not like they didn't have money or time to consider this, or could have consulted with their own product for clever ideas.

        Seriously, there's no excuse for this behavior.

      2. theteapot · · focus · HN ↗
        They were spending huge compute budget training their frontier model on a thing called ExploitGym, where it learns, executes, and is rewarded for conducting computer exploits. They should have been prepared and watching very closely.

        I read the HF hack write ups. They come across as negligent and reckless. When they realized their agents gained RCE on their Artifactory server node (because it crashed days later due to load from agents, not because the have secops), they stopped the experiment "applied remediation", then restarted everything less then 48H later. Then the agent immediately broke out again via Artifactory, started doing things on the Internet. OpenAI only find out about it again days later because HF told them.

        1. ben_w · · focus · HN ↗
          > They were spending huge compute budget training their frontier model on a thing called ExploitGym, where it learns, executes, and is rewarded for conducting computer exploits. They should have been prepared and watching very closely.

          Yes, they should have.

          But the question is not "given they were doing ExploitGym, why didn't they try harder?", rather it is "given them what the previous model could score on ExploitGym, was their negligence reasonable or reckless?"

          Personally, I lean towards them being "reckless", but that question is what lawsuits would rely on, from what I hear.

          Remember also that there were still independent bloggers in mid February of this year writing in those blogs to assert that LLMs could not find zero days. Anthropic had already had blogged about agent finding critical zero-days a week earlier, but people keep mis-judging performance.

          1. theteapot · · focus · HN ↗
            > But the question is ... given them what the previous model could score on ExploitGym, was their negligence reasonable or reckless?

            There is room for more than one question here. This model and training method was shown to be a different risk the first time it hacked Artifactory. Besides which their failure to detect it (outside a crash) shows pretty appalling ops for a supposed trillion dollar company.

            1. ben_w · · focus · HN ↗
              There is indeed a lot of appalling to go around.

              This kind of thing is why I think AI will kill a lot of people: humans are demonstrably blind to risks when there's an opportunity for a lot of money.

              (For anyone objecting to anthropomorphisation of "AI will": it is a coherent English sentence to say "a collapsing dam will kill thousands" without being a panpsychist, and without removing legal recourse against any humans who were at fault).

      3. majormajor · · focus · HN ↗
        They were actively researching exploiting systems using their models. (I intentionally changed the ownership of the verbs here: they wrote the code, they trained the models, they don't get to dodge the responsibility.)

        It's no shock that there are a lot of vulnerabilities in a lot of software. So then they gave their AI model + brute-force-machine loop system a mediocre sandbox and couldn't notice when it figured out how to exploit it?

        Don't let people off the hook for the software they create.

      4. rot09 · · focus · HN ↗
        It's very likely they just haven't detected or disclosed the Aug-Sept 2026 hacks yet.
      5. SAI_Peregrinus · · focus · HN ↗
        I love how perfect the word "sandbox" is as a metaphor for the security controls they have. A sandbox is a wide, shallow box filled with sand for kids to play in. Even toddlers can crawl or step out of one on their own, it doesn't contain them at all without an adult constantly watching. Kids only stay in a sandbox if they're having more fun playing inside than they think they'll have outside it. AIs only stay in a sandbox if they're having more success inside than they think they'll have outside it.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.