‹ BackHN Continuity

Thread

OpenAI agent hacked Australian government website, PM says

256 points · 198 comments · rudy6912

  1. bplatta · · focus · HN ↗
    I'm a little confused as to why these agents are capable of escaping containment. Can someone who has more understanding (or a better guess) of the harness they are running with shed some light?

    To establish the premise: as someone who has a fairly good understanding of the token completion mechanics of an LLM, these agents are completion token calls in a loop, producing a "do this now" request which the harness then runs with some standard "call this function" code.

    If these agents are enabled with explicit network enabled tools, its trivial to monitor their inputs/outputs. If they are not, you can still lock down network egress on a machine. If _some_ network egress is necessary you can still do network traffic monitoring. I don't see how they couldn't implement some level of monitoring where big red lights start flashing when, say, their eval system was contacting a domain/IP located in Australia, and further categorize that domain as government owned. This all seems very doable - am I mistaken?

    And you're telling me all of these companies are failing to do this? Is my understanding naive in some way? This is assuming some good faith of course, I can easily speculate as to the political and corporate incentive. But it seems to me quite risky/negligent.

    Currently, my conclusion is that its just (silly until proven wildly dangerous) negligence with the small side effect of being potentially good for business. And potentially company Foobook is then incentivized to get in on the news cycle for marketing purposes and basically guarantees an agent will do something of the sort by running some harness that allows the behavior quite trivially.

    My naiveté extends to why there is such concern with "losing control of agents" when the above measures seem so doable. It might take a law but it seems doable.

    1. numeri · · focus · HN ↗
      The models in these breaches have already been caught using proxies to get to servers outside of the approved list.

      They've also hacked third party machines and used them to launch attacks on further services.

    2. PunchyHamster · · focus · HN ↗
      > I'm a little confused as to why these agents are capable of escaping containment. Can someone who has more understanding (or a better guess) of the harness they are running with shed some light?

      Probably because containment was written by LLM just to check a box of "we have it contained". Or, just laziness

      With today's internet, there is a good chance just allowing access to a site isn't enough, you might need to give access to 3rd party URLs the site uses.

      ...and if URL for service site uses is same (there is no bucket prefix like for say S3), giving access to service X used by site Y gives access to more than site strictly needs

      ...and if they use cloud stuff people get lazy and just do "allow it entirety of S3 access" vs whitelisting per bucket.

      URL whitelisting wasn't great 20 years ago, now it is just pretty bad for anything cloud based

    3. zobzu · · focus · HN ↗
      they're not really contained and they usually run fully unattended, with a start prompt.

      The first one that was used to market anthropic models was run by a company that called them sandboxed with no Internet access and of course it was disclosed later that they in fact did have internet access, and they won't disclose the prompts.

      insert meme of kid putting a stick in their bike front wheel here... that's how many use LLMs today. it will get worse.

    4. chrisjj · · focus · HN ↗
      > This all seems very doable - am I mistaken?

      Doable by competents? Yes.

      Doable by an outfit that's handed most of its coding to stochastic parrots? Not much chance.

    5. no_multitudes · · focus · HN ↗
      > If these agents are enabled with explicit network enabled tools, its trivial to monitor their inputs/outputs.

      For whatever reason, the AI companies are (or at least were) not doing this kind of classification online during their testing runs, and instead just checking transcripts after the fact. This is more clear in the Anthropic reports about their incidents, for example:

      "The earliest incidents date to April ... We began our transcript review on Thursday, July 23, and stopped all cyber evaluations the same day after identifying transcripts where Claude may have accessed the internet"[1]

      I agree that it is crazy and negligent! I don't think it's good for their business, though -- who wants to use a model that will just cheat instead of doing the job you asked for?

      > My naiveté extends to why there is such concern with "losing control of agents" when the above measures seem so doable. It might take a law but it seems doable.

      At some point, if you are making an LLM in order to use it for useful work, it really benefits you to give it broad network egress.

      [1] <a href="https:&#x2F;&#x2F;www.anthropic.com&#x2F;news&#x2F;investigating-incidents-cybersecurity-evals" rel="nofollow">https:&#x2F;&#x2F;www.anthropic.com&#x2F;news&#x2F;investigating-incidents-cyber...

    6. deaton · · focus · HN ↗
      They&#x27;re capable of escaping containment because headlines are good for their investors. And we don&#x27;t get to see the prompts but who knows what kind of nudging is being done to get these results.
    7. redox99 · · focus · HN ↗
      A lot of it is that the containment was vibecoded (and using older models than what the currently have).
    8. ceroxylon · · focus · HN ↗
      As a network guy, this point has always made me the itchiest. How is it that I can sandbox agents better than a multi-billion dollar company?

      The stakes were that high and they weren&#x27;t running constant PCAPs with DPI and a bunch of alerts? All of that activity to the package manager would have immediately set off multiple alarms in my lab if I was trying to keep things contained. Something doesn&#x27;t add up.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.