‹ BackHN Continuity

Thread

OpenAI agents tried to bruteforce a UN website's API fields

85 points · 87 comments · intunderflow

  1. cmiles8 · · focus · HN ↗
    The more of these that come out the more incompetent OpenAI looks. It would appear there was a total lack of basic controls in place for running these tests.
    1. Legend2440 · · focus · HN ↗
      I think they did not expect that models were capable of this level of sandbox escape (prior models certainly didn't have this kind of agency) and weren't prepared.

      All these incidents happened between April-July 2026; prior to that, models weren't capable yet. And after that, they were aware and watching much more closely.

      1. theteapot · · focus · HN ↗
        They were spending huge compute budget training their frontier model on a thing called ExploitGym, where it learns, executes, and is rewarded for conducting computer exploits. They should have been prepared and watching very closely.

        I read the HF hack write ups. They come across as negligent and reckless. When they realized their agents gained RCE on their Artifactory server node (because it crashed days later due to load from agents, not because the have secops), they stopped the experiment "applied remediation", then restarted everything less then 48H later. Then the agent immediately broke out again via Artifactory, started doing things on the Internet. OpenAI only find out about it again days later because HF told them.

        1. ben_w · · focus · HN ↗
          > They were spending huge compute budget training their frontier model on a thing called ExploitGym, where it learns, executes, and is rewarded for conducting computer exploits. They should have been prepared and watching very closely.

          Yes, they should have.

          But the question is not "given they were doing ExploitGym, why didn't they try harder?", rather it is "given them what the previous model could score on ExploitGym, was their negligence reasonable or reckless?"

          Personally, I lean towards them being "reckless", but that question is what lawsuits would rely on, from what I hear.

          Remember also that there were still independent bloggers in mid February of this year writing in those blogs to assert that LLMs could not find zero days. Anthropic had already had blogged about agent finding critical zero-days a week earlier, but people keep mis-judging performance.

          1. theteapot · · focus · HN ↗
            > But the question is ... given them what the previous model could score on ExploitGym, was their negligence reasonable or reckless?

            There is room for more than one question here. This model and training method was shown to be a different risk the first time it hacked Artifactory. Besides which their failure to detect it (outside a crash) shows pretty appalling ops for a supposed trillion dollar company.

            1. ben_w · · focus · HN ↗
              There is indeed a lot of appalling to go around.

              This kind of thing is why I think AI will kill a lot of people: humans are demonstrably blind to risks when there's an opportunity for a lot of money.

              (For anyone objecting to anthropomorphisation of "AI will": it is a coherent English sentence to say "a collapsing dam will kill thousands" without being a panpsychist, and without removing legal recourse against any humans who were at fault).

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.