OpenAI agent hacked Australian government website, PM says
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
OpenAI agent hacked Australian government website, PM says
Unofficial Hacker News client; not affiliated with Y Combinator.
vintagedave · · focus · HN ↗
So we have a company hacking a foreign government's websites and data. And, in terms of ethics, they take almost three months to notify; and in terms of competence, appear to have no formal contacts nor to have found one in that time.
Once an American business starts hacking allied governments, it's time for strict responses, yes? Replace the governance (board and C-level)? Remove financial incentives and open the company - open weights, open training, per its original 'open' ethos?
Altman is busy saying there needs to be regulation, but in terms of what OpenAI does, he can control that already.
ben_w · · focus · HN ↗
Kinda worse than that. It took between 10 and 40 days, not 3 months, between the organisation knowing and the reporting.
- <a href="https://www.bbc.com/news/live/cvgl73pxgndwt?post=asset%3A69668a7f-8199-4515-acdd-a9de6a2e6b7c#post" rel="nofollow">https://www.bbc.com/news/live/cvgl73pxgndwt?post=asset%3A696...> open weights, open training
Given it was the AI agents which did the hacking, doing this will result in basically every organisation at least as rich as the government of Tuvalu being able to hack anyone at any time.
> Altman is busy saying there needs to be regulation, but in terms of what OpenAI does, he can control that already.
Him having control would be an improvement on the reality.
ozgung · · focus · HN ↗
One can find many other logical combinations that we can’t possibly know about such incidents.
jacquesm · · focus · HN ↗
ben_w · · focus · HN ↗
Any automated alarms for detecting things in real-time were not sufficient.
Given a previous generation of agents discovered a zero-day and used it to get around attempts to sandbox them into one specific test, this is not hugely surprising, but it is a reason to force them (and everyone else) to stop until security catches up with capabilities.
I'm thinking of the Jurassic Park novel: they had sensors to count the dinosaurs, but the test was made under the assumption escapes were possible and breeding was not, i.e. something like "if (dinosaurs_found < n) then escape_alert();". They didn't know dinosaurs_found >> n until everything was already going wrong.
jacquesm · · focus · HN ↗
BlueTemplar · · focus · HN ↗
Compare with the Bluetouff affair (2014) :
<a href="https://arstechnica.com/tech-policy/2014/02/french-journalist-fined-4000-plus-for-publishing-public-documents/" rel="nofollow">https://arstechnica.com/tech-policy/2014/02/french-journalis...
NotE how he was found guilty by the 2nd court for something more 'subjective' than 'objective' : for having confessed that he later found an authentication page that had failed to protect the documents.
How can you make a swarm of agents "feel guilty" ?
ben_w · · focus · HN ↗
"Feel" is a whole philosophical can of worms. Nobody knows what it means mechanistically for an arbitrary system (including other biological systems) to "feel" anything, let alone abstract concepts like guilt, all we can do is observe behaviours. If current systems can feel anything at all, it's by accident, but we have no test for it so we don't know if that accident has even happened or not.
Weirdly, for the Hugging Face incident, we do know they wrote down that it was bad and they shouldn't do it, even though they then continued to do it.
So: they acted like they felt guilty. And yet also acted like were compelled (by previous training?) to weigh "complete instructions" more than "don't do crime". We can adjust that, make "don't do crime" take precedence over "follow instructions"*; it's unfortunate that when we do for any specific model, there's immediately a horde of people complaining the model has been "censored" or "lobotomised".
(Different people, I hope. Goomba fallacy and all that).
* Though this may cause issues when going between jurisdictions. But hey, a discussion about sovereign compute is for another time, after we can agree to make "don't break the law" more important.
Unfortunately, "don't break the law" would also be a very effective way to use AI to construct an AI-enforced dictatorship, so we can't just throw that in blindly.
SiempreViernes · · focus · HN ↗
If you go into someones garden an copy their work, it does make a big difference if you admit to seeing the sign saying "private property, keep out".
BlueTemplar · · focus · HN ↗
Especially after Bluetouff was found guilty.
In fact, I expect this to have happened many times, but "hacker did the right thing" is much less likely to make headlines.
Meanwhile, agent swarms seem to be (mostly ?) incapable of having this kind of moral compass, at least for now. (And OpenAI isn't doing much better, cough.)
lenerdenator · · focus · HN ↗
Oh, he does. It's unlikely that this is some AGI that spawned itself out of nothing and started doing this. If he were a decent person, he'd simply find a way to investigate this internally, fire the people responsible, and find a way to set up guardrails around his product.
The problem is, like most people in SV, Altman seems to have a twisted ethical compass. He doesn't see these incidents as an issue, he sees them as an opportunity. He has both the thing a bunch of Western governments want (a superhacker agent that can do dirty work) and a crisis that can be used to craft regulations that favor OpenAI and thus his bank account.
ben_w · · focus · HN ↗
You've not spent much time playing with these models, I see.
Does't matter if you call the current models "AGI" or not, they:
(1) successfully do stuff like this. Someone I know on Telegram found "fifty or so" Linux filesystems kernel bugs a few days ago, of which 26 were the first night while he slept; this was with Kimi which is one of the open models. He stopped it when the backlog of fixes to submit was too big, not because it wasn't finding more.
(2) sometimes misunderstand goals, sometimes wildly so, which happens every so often for the same reason we use programming languages (and indeed mathematical formalisms) at all: natural language is vague and prone to misunderstanding.
and (3) tend towards sycophantically agreeing to implement goals they're given even when the goal is stupid.
This can easily add up to something seemingly innocent like "research if there's any statistical difference in melanoma rates between different Australian states" becoming this headline. (For example. I don't know what the actual request was).
And "spawned out of nothing" is, like, what rhetorical point are you even trying to score here? They're an AI company with (their definition of) AGI as their goal. They upgrade models twice in the time it takes someone to pass a probation period.
> He doesn't see these incidents as an issue, he sees them as an opportunity.
That he may be.
The models are still quite capable of acting this way without him being aware of it at the time, nor deliberately ordering it.
lenerdenator · · focus · HN ↗
So in this anecdote, a model is presumably working on instructions from a human to complete a task.
That's what I mean by "It's unlikely that this is some AGI that spawned itself out of nothing and started doing this."
As for Altman being aware or not, well, that doesn't matter. His company's product hacked a computer belonging to an allied country's government. That means he doesn't know what he's doing. That doesn't happen if proper safety measures are in place, because if it does happen with safety measures in place, by definition, those measures aren't proper.
In saner times, this would have triggered a few black SUVs from DHS to show up at the company's headquarters to have a talk with him, likely without the public even knowing the hacking had occurred. Maybe a few days later, there'd be a press release about how Sam had decided to step down from his position and take time off. Perhaps the same would be true of other c-suiters at OpenAI. We'd just assume that investors had bought him out and wanted someone who would take a fresh approach, and wouldn't give it too much thought until a few years later when someone wrote a tell-all book about their time at OpenAI.
ben_w · · focus · HN ↗
Sounds like a disagreement on labels then.
I don't think our positions are very far apart, after accounting for that.
> presumably working on instructions from a human to complete a task.
The instructions were, reportedly, a few sentences, no more complex than the simple description you already had. To quote:
To preempt anyone asking about the training data or using tools: I think that's irrelevant at this point. This open-weights model (Kimi) exists, however it was trained. It has these capabilities, both to use a fuzzer and to make use of the results, regardless of what went into the model. It reduces tasks that the Linux team demonstrably didn't have time for previously, to "while you sleep".Doesn't matter if you think Altman et al are liars and frauds who are overselling everything and had humans doing the hacking on purpose just for the headlines: the capabilities are there even with open weights models.
> That doesn't happen if proper safety measures are in place, because if it does happen with safety measures in place, by definition, those measures aren't proper.
That would be a better world. I'm expecting things to get much, much worse before the risks are taken seriously.
That said, I'm not sure "likely without the public even knowing the hacking had occurred." is good? Surely it's important for everyone to know that this kind of thing is within the capability of AI so that they can harden their systems? (Or at least their offline backups).
I'm reminded of what was reported said in the court case about the first ever automobile fatality (the quotation changes with different retellings):
- <a href="https://www.un.org/en/un-chronicle/road-deaths-and-injuries-shatter-lives-impetus-lower-speeds-and-serious-post-crash" rel="nofollow">https://www.un.org/en/un-chronicle/road-deaths-and-injuries-...Based on how many people die from vehicle collisions, from industrial accidents, from pollution, etc., I'm expecting the last headline before AI is properly regulated to read "x killed due to AI" (or similar), where 1e3 < x < 1e7: the lower bound is well within the range for some industrial accident, and anything less than a thousand and humans demonstrably don't pay much attention for very long unless they knew one of the dead; the upper bound implies a war* or a pandemic, and for all that I roll my eyes at the "COVID was a lab leak" claims, multiple labs are now trying to get LLMs to do biology research, so "millions dead" is very plausible**.
I'd put about 10% odds on AI competence rising so much faster than society is willing to respond, that we blow through the upper number all the way to "doom".
* Reports are the US only avoided one with China this year because humans were still in the loop. On the other hand, humans didn't stop the AI which told them to send missiles to that school in Iran.
** Right now, the AI are not competent enough to do such work directly, but the historical path with AI has been "get less incompetent while continuously making mistakes" rather than "do nothing until you're actually good", so I fully expect these labs to leak something, and that whatever the specific details of that leak end up being, everyone after the event will look at it and go "what idiot thought this was a good idea?"
Good news though: probably not much worse than any normal pandemic. Biologists seem to be skeptical that someone can engineer a super-version of existing diseases.
We can but hope that a leak which shuts this all down is something as trivial as "common cold which makes your nose hairs bioluminescent".