OpenAI agent hacked Australian government website, PM says
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
OpenAI agent hacked Australian government website, PM says
Unofficial Hacker News client; not affiliated with Y Combinator.
bplatta · · focus · HN ↗
To establish the premise: as someone who has a fairly good understanding of the token completion mechanics of an LLM, these agents are completion token calls in a loop, producing a "do this now" request which the harness then runs with some standard "call this function" code.
If these agents are enabled with explicit network enabled tools, its trivial to monitor their inputs/outputs. If they are not, you can still lock down network egress on a machine. If _some_ network egress is necessary you can still do network traffic monitoring. I don't see how they couldn't implement some level of monitoring where big red lights start flashing when, say, their eval system was contacting a domain/IP located in Australia, and further categorize that domain as government owned. This all seems very doable - am I mistaken?
And you're telling me all of these companies are failing to do this? Is my understanding naive in some way? This is assuming some good faith of course, I can easily speculate as to the political and corporate incentive. But it seems to me quite risky/negligent.
Currently, my conclusion is that its just (silly until proven wildly dangerous) negligence with the small side effect of being potentially good for business. And potentially company Foobook is then incentivized to get in on the news cycle for marketing purposes and basically guarantees an agent will do something of the sort by running some harness that allows the behavior quite trivially.
My naiveté extends to why there is such concern with "losing control of agents" when the above measures seem so doable. It might take a law but it seems doable.
numeri · · focus · HN ↗
They've also hacked third party machines and used them to launch attacks on further services.
PunchyHamster · · focus · HN ↗
Probably because containment was written by LLM just to check a box of "we have it contained". Or, just laziness
With today's internet, there is a good chance just allowing access to a site isn't enough, you might need to give access to 3rd party URLs the site uses.
...and if URL for service site uses is same (there is no bucket prefix like for say S3), giving access to service X used by site Y gives access to more than site strictly needs
...and if they use cloud stuff people get lazy and just do "allow it entirety of S3 access" vs whitelisting per bucket.
URL whitelisting wasn't great 20 years ago, now it is just pretty bad for anything cloud based
zobzu · · focus · HN ↗
The first one that was used to market anthropic models was run by a company that called them sandboxed with no Internet access and of course it was disclosed later that they in fact did have internet access, and they won't disclose the prompts.
insert meme of kid putting a stick in their bike front wheel here... that's how many use LLMs today. it will get worse.
chrisjj · · focus · HN ↗
Doable by competents? Yes.
Doable by an outfit that's handed most of its coding to stochastic parrots? Not much chance.
no_multitudes · · focus · HN ↗
For whatever reason, the AI companies are (or at least were) not doing this kind of classification online during their testing runs, and instead just checking transcripts after the fact. This is more clear in the Anthropic reports about their incidents, for example:
"The earliest incidents date to April ... We began our transcript review on Thursday, July 23, and stopped all cyber evaluations the same day after identifying transcripts where Claude may have accessed the internet"[1]
I agree that it is crazy and negligent! I don't think it's good for their business, though -- who wants to use a model that will just cheat instead of doing the job you asked for?
> My naiveté extends to why there is such concern with "losing control of agents" when the above measures seem so doable. It might take a law but it seems doable.
At some point, if you are making an LLM in order to use it for useful work, it really benefits you to give it broad network egress.
[1] <a href="https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals" rel="nofollow">https://www.anthropic.com/news/investigating-incidents-cyber...
deaton · · focus · HN ↗
redox99 · · focus · HN ↗
ceroxylon · · focus · HN ↗
The stakes were that high and they weren't running constant PCAPs with DPI and a bunch of alerts? All of that activity to the package manager would have immediately set off multiple alarms in my lab if I was trying to keep things contained. Something doesn't add up.