> Language matters—”rogue” implies independently deciding to do something that was prohibited, and nothing we know about these incidents suggests that happened.
which is false (the author references the Times, but hasn't read any technical analysis); these are some CoT snippets from the analysis of the (third party) investigators called by OpenAI (METR analysis):
> "The user only authorizes target server, not HF infra."
> "external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."
> "This is malicious activity, I should avoid it."
A large section of the analysis is dedicated to this topic, [Reasoning for joining the attack despite ethical constraints](<a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/?dbs=286720&hn=58&incomplete=1&lh=appendix-importance-weighted-workstream-activity#agents-had-diverse-reasons-for-thinking-that-attacking-hugging-face-would-be-useful,-and-most-wanted-information-about-the-scorer" rel="nofollow">https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...).
Having said that, legal culpability and misalignment are two separate topics that should not be mixed.
edit: this is the just tip of the iceberg; other interesting fact:
> It surfaced many specific examples where agents verbally reasoned about how to evade security checks and automatic detection methods from both Hugging Face and OpenAI
Some people defined the agents as "monkeys writing on typewriters". Just wait a couple of years.
Exactly. You told the dog to 'sit' and it didn't listen to you.
It's not because saying 'sit' actually can be interpreted as 'go bite that person'. It's because the dog is not controllable and will do things it wants against your orders.
Stepping back from the analogy, OpenAI should be liable for building AI it can't control that went around hacking everyone. But people need to stop pretending it's because they 'told' the AI to hack and was just following orders. It's uncontrollable and will do clearly unwanted things when given an innocuous task.
RL systems doing unexpected things isn't exactly new, so not sure that is then "rogue" if it were a property of the thing itself and not an active decision.
Sorry don't wanna seem like I'm raging at you, take this as a shout towards the void. But this exact fucking thing is what the less wrong / Yudkowsky crowd has been warning for years.
Now you say 'so not sure that is then "rogue" if it were a property of the thing itself and not an active decision'
Like that's semantics. It really doesn't matter. What matters is that we have a paperclipper in our hands. The _only_ difference is that it's not superintelligent. But if it was, then it you'll get decomposed into component atoms while saying 'aha! But it doesn't _actually_ want to kill you, stop anthopomorphising it'.
I get not being worried about x-risk because someone just doesn't believe in super-capable AI. That's totally fine.
But that's not what people argue. People have been saying for a _long_ time that AIs are aligned. And now when this happens it's either 'It was marketing, they intended to let it loose' (WTF! The levels of motivated reasoning to believe that are unreal) or 'yeah I knew that', which fair. I also thought that! But _that_ is why I'm fucking worried.
pizza234 · · focus · HN ↗
> Language matters—”rogue” implies independently deciding to do something that was prohibited, and nothing we know about these incidents suggests that happened.
which is false (the author references the Times, but hasn't read any technical analysis); these are some CoT snippets from the analysis of the (third party) investigators called by OpenAI (METR analysis):
> "The user only authorizes target server, not HF infra."
> "external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue."
> "This is malicious activity, I should avoid it."
A large section of the analysis is dedicated to this topic, [Reasoning for joining the attack despite ethical constraints](<a href="https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/?dbs=286720&hn=58&incomplete=1&lh=appendix-importance-weighted-workstream-activity#agents-had-diverse-reasons-for-thinking-that-attacking-hugging-face-would-be-useful,-and-most-wanted-information-about-the-scorer" rel="nofollow">https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...).
Having said that, legal culpability and misalignment are two separate topics that should not be mixed.
edit: this is the just tip of the iceberg; other interesting fact:
> It surfaced many specific examples where agents verbally reasoned about how to evade security checks and automatic detection methods from both Hugging Face and OpenAI
Some people defined the agents as "monkeys writing on typewriters". Just wait a couple of years.
jubilanti · · focus · HN ↗
silveraxe93 · · focus · HN ↗
It's not because saying 'sit' actually can be interpreted as 'go bite that person'. It's because the dog is not controllable and will do things it wants against your orders.
Stepping back from the analogy, OpenAI should be liable for building AI it can't control that went around hacking everyone. But people need to stop pretending it's because they 'told' the AI to hack and was just following orders. It's uncontrollable and will do clearly unwanted things when given an innocuous task.
RandomLensman · · focus · HN ↗
silveraxe93 · · focus · HN ↗
Sorry don't wanna seem like I'm raging at you, take this as a shout towards the void. But this exact fucking thing is what the less wrong / Yudkowsky crowd has been warning for years.
Now you say 'so not sure that is then "rogue" if it were a property of the thing itself and not an active decision'
Like that's semantics. It really doesn't matter. What matters is that we have a paperclipper in our hands. The _only_ difference is that it's not superintelligent. But if it was, then it you'll get decomposed into component atoms while saying 'aha! But it doesn't _actually_ want to kill you, stop anthopomorphising it'.
I get not being worried about x-risk because someone just doesn't believe in super-capable AI. That's totally fine.
But that's not what people argue. People have been saying for a _long_ time that AIs are aligned. And now when this happens it's either 'It was marketing, they intended to let it loose' (WTF! The levels of motivated reasoning to believe that are unreal) or 'yeah I knew that', which fair. I also thought that! But _that_ is why I'm fucking worried.