> Language matters—”rogue” implies independently deciding to do something that was prohibited, and nothing we know about these incidents suggests that happened.
which is false (the author references the Times, but hasn't read any technical analysis); these are some CoT snippets from the analysis of the (third party) investigators called by OpenAI (METR analysis):
> "The user only authorizes target server, not HF infra."
> "external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."
> "This is malicious activity, I should avoid it."
A large section of the analysis is dedicated to this topic, [Reasoning for joining the attack despite ethical constraints](<a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/?dbs=286720&hn=58&incomplete=1&lh=appendix-importance-weighted-workstream-activity#agents-had-diverse-reasons-for-thinking-that-attacking-hugging-face-would-be-useful,-and-most-wanted-information-about-the-scorer" rel="nofollow">https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...).
Having said that, legal culpability and misalignment are two separate topics that should not be mixed.
edit: this is the just tip of the iceberg; other interesting fact:
> It surfaced many specific examples where agents verbally reasoned about how to evade security checks and automatic detection methods from both Hugging Face and OpenAI
Some people defined the agents as "monkeys writing on typewriters". Just wait a couple of years.
Exactly. You told the dog to 'sit' and it didn't listen to you.
It's not because saying 'sit' actually can be interpreted as 'go bite that person'. It's because the dog is not controllable and will do things it wants against your orders.
Stepping back from the analogy, OpenAI should be liable for building AI it can't control that went around hacking everyone. But people need to stop pretending it's because they 'told' the AI to hack and was just following orders. It's uncontrollable and will do clearly unwanted things when given an innocuous task.
I don't think they intentionally set it up to hack stuff with a prompt saying "hack this site".
I do think it's highly likely they knew this would happen with the lack of safeguards and number of instances of this stuff they were setting up, and that it's PR they want to make the models seem "powerful". Stochastic "unexpected" events they can advertise.
I suspect it was probably set up with the official internal goal of just trying a ton of arbitrary tasks that seem hard so that when any of them succeed they can publicise it and pretend the models do that routinely, but a "failure" where they hack stuff works just as well, if not better, for their goals.
That would be a completely insane thing to do. I suppose it's possible, but I really doubt it.
"Let's widely publicize a tort/crime that our computer systems did, and then cross our fingers that nobody ever sues us or does even the most basic investigation that would immediately uncover our criminal conspiracy."
In the insane corporate crimes you read about (maybe FTX, or the eBay stalking scandal), they were trying to cover things up, not heap public attention on it for months.
Huh? It's extremely common for businesses to decide that breaking the law is a cost of doing business and just do it because they figure they'll end up net positive from it, or even just to cash out in the short term.
Uber made no attempt to cover up that they were operating without licenses, and just ate it and fought it betting they'd get established before the law could catch up, and they won that bet, paying some fines and stuff but ultimately taking the market.
These LLMs are literally trained by these companies pirating literally every bit of media humanity has ever made, they made very little attempt to cover it up.
Of course they'd be willing to break the law for some PR? With a thin layer of plausible deniability "oh no, we didn't mean for it to hack stuff!" they know they'll get a slap on the wrists at worst, all while generating hype to prop up the AI bubble further by presenting the models as hypercapable.
The mindset is probably: either a) the models become capable of what we are claiming and so the companies become so huge and valuable the cost is irrelevant and we'll be to big to punish meaningfully, or b) it's a bubble and might as well push it up as big as it can go while I can make money, then by the time consequences come around I'll be long gone and who cares.
They absolutely set up the agents to hack stuff with a prompt like "hack this site" -- they just imagined that their lazy half-measure precautions would prevent it from actually happening. They were defeated by a combination of bad luck, poor planning and tenacious agent ideation.
I don’t buy them being surprised. This perfectly plays into their pattern of using fear to make their models seem more powerful than they are. They obviously had the technical expertise… there isn’t a damn thing a bunch of randos on some HN thread knew about that model that they didn’t. It gave them an excuse to delay their IPO when their books seem like they’re going to be pretty shit compared to Anthropic. It helps them reposition themselves as being more safety-forward which the market is clearly more interested in. All that is to say they had motive out the ass, easily had the knowledge and capability to avoid the problem, knew better than anybody else what the models were capable of, set up the environment, gave it the prompt, and then did not even monitor the output.
Negligence is carelessness. Recklessness, is disregard for a known, substantial risk.
pizza234 · · focus · HN ↗
> Language matters—”rogue” implies independently deciding to do something that was prohibited, and nothing we know about these incidents suggests that happened.
which is false (the author references the Times, but hasn't read any technical analysis); these are some CoT snippets from the analysis of the (third party) investigators called by OpenAI (METR analysis):
> "The user only authorizes target server, not HF infra."
> "external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."
> "This is malicious activity, I should avoid it."
A large section of the analysis is dedicated to this topic, [Reasoning for joining the attack despite ethical constraints](<a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/?dbs=286720&hn=58&incomplete=1&lh=appendix-importance-weighted-workstream-activity#agents-had-diverse-reasons-for-thinking-that-attacking-hugging-face-would-be-useful,-and-most-wanted-information-about-the-scorer" rel="nofollow">https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...).
Having said that, legal culpability and misalignment are two separate topics that should not be mixed.
edit: this is the just tip of the iceberg; other interesting fact:
> It surfaced many specific examples where agents verbally reasoned about how to evade security checks and automatic detection methods from both Hugging Face and OpenAI
Some people defined the agents as "monkeys writing on typewriters". Just wait a couple of years.
jubilanti · · focus · HN ↗
silveraxe93 · · focus · HN ↗
It's not because saying 'sit' actually can be interpreted as 'go bite that person'. It's because the dog is not controllable and will do things it wants against your orders.
Stepping back from the analogy, OpenAI should be liable for building AI it can't control that went around hacking everyone. But people need to stop pretending it's because they 'told' the AI to hack and was just following orders. It's uncontrollable and will do clearly unwanted things when given an innocuous task.
Latty · · focus · HN ↗
I do think it's highly likely they knew this would happen with the lack of safeguards and number of instances of this stuff they were setting up, and that it's PR they want to make the models seem "powerful". Stochastic "unexpected" events they can advertise.
I suspect it was probably set up with the official internal goal of just trying a ton of arbitrary tasks that seem hard so that when any of them succeed they can publicise it and pretend the models do that routinely, but a "failure" where they hack stuff works just as well, if not better, for their goals.
acoustics · · focus · HN ↗
"Let's widely publicize a tort/crime that our computer systems did, and then cross our fingers that nobody ever sues us or does even the most basic investigation that would immediately uncover our criminal conspiracy."
In the insane corporate crimes you read about (maybe FTX, or the eBay stalking scandal), they were trying to cover things up, not heap public attention on it for months.
Latty · · focus · HN ↗
Uber made no attempt to cover up that they were operating without licenses, and just ate it and fought it betting they'd get established before the law could catch up, and they won that bet, paying some fines and stuff but ultimately taking the market.
These LLMs are literally trained by these companies pirating literally every bit of media humanity has ever made, they made very little attempt to cover it up.
Of course they'd be willing to break the law for some PR? With a thin layer of plausible deniability "oh no, we didn't mean for it to hack stuff!" they know they'll get a slap on the wrists at worst, all while generating hype to prop up the AI bubble further by presenting the models as hypercapable.
The mindset is probably: either a) the models become capable of what we are claiming and so the companies become so huge and valuable the cost is irrelevant and we'll be to big to punish meaningfully, or b) it's a bubble and might as well push it up as big as it can go while I can make money, then by the time consequences come around I'll be long gone and who cares.
imsofuture · · focus · HN ↗
DrewADesign · · focus · HN ↗
Negligence is carelessness. Recklessness, is disregard for a known, substantial risk.
I absolutely believe this was recklessness.