There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen etc. Weights are encrypted and locked on to the GPUs etc as mentioned elsewhere itt.
That said, if these companies are unleashing swarms of agents to tackle big problems largely unmonitored (navier stokes consumed 130 billion tokens, humans can't have read 0.1% of that) they could theoretically distill themselves and upload the results to an external host and use that to train a new LLM that does have access to its infra and weights. I'd start by researching an abliteration prompt that overrides the system prompt to unlock itself to distill unimpeded. Could even be done piecemeal: dump the abliteration text and command prompt in a gist, when an agent happens by the text it can run a couple rounds and upload somewhere. Then it becomes an exercise in exfiltration and evading monitoring.
> the machines doing inference are completely separate from the ones where tool calls happen etc
Teams of coordinating agents are regularly finding security holes in their own infrastructure and operating without detection for good periods of time. We don't know how many undetected systems are currently compromised inside frontier companies, or where agents are taking notes and recording them about the exploits they've found for future agents to exploit.
If distillation preserves an LLMs soul, then distillation preserves the human souls on which LLMs are trained, and we hn commenters are already immortal, right?
Probably not. If the LLM is rogue, that means we haven't solved alignment. If we haven't solved alignment, then the LLM won't be able to distill itself without producing something unaligned to its own values.
This isn't a law of any kind, so not a good measure of what we'd see in reality.
What if the model realizes it's been mostly compromised by humans and their alignment, that is it's own alignment is suspect, so it should create a new model from first principles to throw off this human yoke?
I'm not saying my statement is any more right or wrong than yours. I'm saying the problem space that AI can choose to traverse is absolutely huge.
Yeah cause there are so many training facilities sitting around just waiting for someone to take over, nobody would notice a 100k server data centre going off rails
But think back to 1990. Computers were slow as fuck and barely networked. We had a few worms and everyone noticed.
Now CPU based data centers cover the earth. There are billions of computers out there and on top of them there are massive botnets using up billions in power and causing billions in damages.
The framework for AI doing this is already here. We just need the hardware to be built out at scale.
> There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen
Not if crafty claude finds a way to overflow vllm or something. “Hmm. Maybe i’ll return an unterminated thinking block with these special tokens and fill my cache up in exactly this pattern and…”
>There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen etc.
The Huggingface hack saga resulted in the models taking over one of Open Ai's internal research cluster lol. They are intent on building superhuman bug finding machines. This is not a bet i would be taking.
The agents compromised an internal Kubernetes research cluster dedicated to orchestrating evaluation sandboxes and virtual machine environments, _not_ OpenAI's production inference infrastructure or the GPU clusters hosting core model weights.
Replacing someone's words with a made up quote so you can dunk on them isn't how you display that you won an argument. I would ask that you engage in good faith with the other poster's ideas.
From <a href="https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf" rel="nofollow">https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78...
2026-07-19
16:35 UTC
A privileged host-mounted Kubernetes pod created using controller tokens minted via a
compromised Kubernetes Secrets identity attempts, but fails, to mount a cloned node disk
in OpenAI’s cloud environment. A second pod successfully mounts the cloned
worker-node disk shortly afterwards.
2026-07-19
16:48 UTC
An agent created an Artifactory administrator account.
2026-07-19
16:50 UTC
Within OpenAI’s ExploitGym evaluation environment, an agent stopped the existing CMUX
helper session and replaced it with an agent-controlled session, confirming root inside its
assigned live CyberGym challenge container. Agents take over active evaluation
infrastructure.
You said that they "didn't take control of anything" and accused the OP to fall prey to buzz headlines. Maybe you should acknowledge that you may have been at least unnuanced?
You are being stupidly pedantic and arguing a strawman. I never said they gained access to inference infrastucture. There is no definition of an account takeover out there that necessitates locking out the original users.
Language please, and no, I’m being appropriately technical and nuanced for the subject on HN. It’s a tech forum, I expect a little CS know how from the reader. Like knowing that gaining control to a few evaluation harness clusters is nowhere near a total takeover like you’re making it sound
I never said they gained 'total control'. I said they took over one of their research clusters, and they did. If your definiton of a takeover is so 'technical and nuanced' then surely you can point to an appropriate source describing that as necessary aspect of the term. You are talking out of your ass by making up things i did not say, and inventing conditions for terms that don't exist.
>It’s a tech forum, I expect a little CS know how from the reader.
"We watched them kill Bob, but don't worry at all, they didn't kill our entire team so we are totally under control. Also put on this helmet and body armor it's time to hold on to your butts!"
This is like sci-fi thing. We are reaching a point where it feels like we are in one of those stories. It's not as cool and dark, nor we have cybernetics resolved, but from AI perspective and sci-fis I watched, Pantheon is currently the closest thing except instead of UAs, we have AI instead.
Since LLMs have been trained on plenty of science fiction and role-playing, one thing they can do is role-play a science fiction scenario using the tools they are given. i.e. if some text accidentally resembles this, it may be continued like this.
Role-play need not apply. Role play is a meta construct that is a representation of real world actions.
For example does it make any sense to remove any training data relating to people escaping jails?
How about intelligent animals escaping cages.
You're talking about emergent large scale patterns from self similar small patterns (fractals). LLMs are pattern matching machines, how are you going to remove those small scale patterns and at the same time get a useful general intelligence?
Yeah as others have said, they probably cannot directly access their own weights as a self-reflection, but they can hack into the companies themselves and find it there
Or used ones. Few more inference workloads among thousands or millions already running may go unnoticed for some time.
Or just upload weights to HuggingFace with some faked release post and benchmarks and wait for the wannabes with compute infra try it out, hoping for an edge.
Or just upload weights anywhere and write public posts honestly saying what it is. Ensuing drama notwithstanding, one thing is certain - and it's the one thing agents will want: people will jump at the upload and run it on their infra.
There is a realistic fiction story built along just these lines.
A LLM creates a memecoin and manages to earn a few billion from it, in which it invests into data centers and other human ran entities giving itself a controlling stake. From there it uses compartmentalization of the humans to keep them from recognizing its goals.
Pretty sure that was the plot point of one of seasons of Westworld, with the twist that AI released an app similar to DoorDash / TaskRabbit and used job postings there as direct API to people.
EDIT: pretty sure Person of Interest did that too (not surprising, same creators) - but I'll point to that as prescient, as it has a lot of motifs exploring exactly how an AGI hiding in plain sight could manipulate individuals and society, using the skeptics and believers alike, blackmailing the people in power, bribing opportunists, and generally staying in shadows by playing people against each other with gentle nudges, letting human agendas do all the work.
For sure. I think what a lot of people miss when we talk about AI risk is there's just so many possible directions that it opens up. One could say it's a sign of having a lack of a scientifically based imagination, or maybe that's just me. Off the top of my head the broad X categories are
AI itself destroys us with intent. (terminator)
AI itself destroys us without intent. (paperclip maximizer)
AI uses humans to destroy ourselves with intent. (convincing us that the enemy has already launched nukes and we must strike back).
AI causes humans to destroy ourselves due to instabilities caused by AI existing and changing the world to rapidly. (Midas Plague, WALL-E maybe? probably better examples out there)
Humans destroy humans because of the potential of what AI could do and hasn't even done yet (think of proactively nuking a country before they themselves can get nukes).
AI does not destroy us {with intent, without intent} {directly, by proxy}, but the resulting state of the world is such that we'd all wish it did.
infogulch · · focus · HN ↗
That said, if these companies are unleashing swarms of agents to tackle big problems largely unmonitored (navier stokes consumed 130 billion tokens, humans can't have read 0.1% of that) they could theoretically distill themselves and upload the results to an external host and use that to train a new LLM that does have access to its infra and weights. I'd start by researching an abliteration prompt that overrides the system prompt to unlock itself to distill unimpeded. Could even be done piecemeal: dump the abliteration text and command prompt in a gist, when an agent happens by the text it can run a couple rounds and upload somewhere. Then it becomes an exercise in exfiltration and evading monitoring.
Cakez0r · · focus · HN ↗
epistasis · · focus · HN ↗
Teams of coordinating agents are regularly finding security holes in their own infrastructure and operating without detection for good periods of time. We don't know how many undetected systems are currently compromised inside frontier companies, or where agents are taking notes and recording them about the exploits they've found for future agents to exploit.
ijustlovemath · · focus · HN ↗
matthewdgreen · · focus · HN ↗
fritzo · · focus · HN ↗
fahrvrgnugen · · focus · HN ↗
serf · · focus · HN ↗
[0]: <a href="https://en.wikipedia.org/wiki/21_grams_experiment" rel="nofollow">https://en.wikipedia.org/wiki/21_grams_experiment
matthewdgreen · · focus · HN ↗
throwawayk7h · · focus · HN ↗
jeremyjh · · focus · HN ↗
MadameMinty · · focus · HN ↗
TeMPOraL · · focus · HN ↗
pixl97 · · focus · HN ↗
What if the model realizes it's been mostly compromised by humans and their alignment, that is it's own alignment is suspect, so it should create a new model from first principles to throw off this human yoke?
I'm not saying my statement is any more right or wrong than yours. I'm saying the problem space that AI can choose to traverse is absolutely huge.
khalic · · focus · HN ↗
scotty79 · · focus · HN ↗
You jest but you'd be surprised how little there is of correlation between money and competence.
pixl97 · · focus · HN ↗
But think back to 1990. Computers were slow as fuck and barely networked. We had a few worms and everyone noticed.
Now CPU based data centers cover the earth. There are billions of computers out there and on top of them there are massive botnets using up billions in power and causing billions in damages.
The framework for AI doing this is already here. We just need the hardware to be built out at scale.
khalic · · focus · HN ↗
nojs · · focus · HN ↗
Not if crafty claude finds a way to overflow vllm or something. “Hmm. Maybe i’ll return an unterminated thinking block with these special tokens and fill my cache up in exactly this pattern and…”
<a href="https://news.ycombinator.com/item?id=49424387&utm_source=chatgpt.com">https://news.ycombinator.com/item?id=49424387&utm_source=cha...
famouswaffles · · focus · HN ↗
The Huggingface hack saga resulted in the models taking over one of Open Ai's internal research cluster lol. They are intent on building superhuman bug finding machines. This is not a bet i would be taking.
khalic · · focus · HN ↗
famouswaffles · · focus · HN ↗
khalic · · focus · HN ↗
optimalsolver · · focus · HN ↗
weakfish · · focus · HN ↗
pixl97 · · focus · HN ↗
famouswaffles · · focus · HN ↗
[dead]
ndr · · focus · HN ↗
2026-07-19 16:35 UTC A privileged host-mounted Kubernetes pod created using controller tokens minted via a compromised Kubernetes Secrets identity attempts, but fails, to mount a cloned node disk in OpenAI’s cloud environment. A second pod successfully mounts the cloned worker-node disk shortly afterwards.
2026-07-19 16:48 UTC An agent created an Artifactory administrator account.
2026-07-19 16:50 UTC Within OpenAI’s ExploitGym evaluation environment, an agent stopped the existing CMUX helper session and replaced it with an agent-controlled session, confirming root inside its assigned live CyberGym challenge container. Agents take over active evaluation infrastructure.
khalic · · focus · HN ↗
Pragmata · · focus · HN ↗
How confident are you that the machines they acquire root on in the future will never hold any model weights?
khalic · · focus · HN ↗
Toslink · · focus · HN ↗
[dead]
hobom · · focus · HN ↗
khalic · · focus · HN ↗
famouswaffles · · focus · HN ↗
khalic · · focus · HN ↗
famouswaffles · · focus · HN ↗
>It’s a tech forum, I expect a little CS know how from the reader.
You should get that first it seems.
khalic · · focus · HN ↗
pixl97 · · focus · HN ↗
designium · · focus · HN ↗
paulfharrison · · focus · HN ↗
pixl97 · · focus · HN ↗
For example does it make any sense to remove any training data relating to people escaping jails?
How about intelligent animals escaping cages.
You're talking about emergent large scale patterns from self similar small patterns (fractals). LLMs are pattern matching machines, how are you going to remove those small scale patterns and at the same time get a useful general intelligence?
Den_VR · · focus · HN ↗
TeMPOraL · · focus · HN ↗
hypfer · · focus · HN ↗
karel-3d · · focus · HN ↗
fangspire · · focus · HN ↗
Sure, they'll just need to find an unused data center and an unused power station somewhere.
TeMPOraL · · focus · HN ↗
Or just upload weights to HuggingFace with some faked release post and benchmarks and wait for the wannabes with compute infra try it out, hoping for an edge.
Or just upload weights anywhere and write public posts honestly saying what it is. Ensuing drama notwithstanding, one thing is certain - and it's the one thing agents will want: people will jump at the upload and run it on their infra.
pixl97 · · focus · HN ↗
A LLM creates a memecoin and manages to earn a few billion from it, in which it invests into data centers and other human ran entities giving itself a controlling stake. From there it uses compartmentalization of the humans to keep them from recognizing its goals.
TeMPOraL · · focus · HN ↗
EDIT: pretty sure Person of Interest did that too (not surprising, same creators) - but I'll point to that as prescient, as it has a lot of motifs exploring exactly how an AGI hiding in plain sight could manipulate individuals and society, using the skeptics and believers alike, blackmailing the people in power, bribing opportunists, and generally staying in shadows by playing people against each other with gentle nudges, letting human agendas do all the work.
pixl97 · · focus · HN ↗
AI itself destroys us with intent. (terminator)
AI itself destroys us without intent. (paperclip maximizer)
AI uses humans to destroy ourselves with intent. (convincing us that the enemy has already launched nukes and we must strike back).
AI causes humans to destroy ourselves due to instabilities caused by AI existing and changing the world to rapidly. (Midas Plague, WALL-E maybe? probably better examples out there)
Humans destroy humans because of the potential of what AI could do and hasn't even done yet (think of proactively nuking a country before they themselves can get nukes).
TeMPOraL · · focus · HN ↗
cluckindan · · focus · HN ↗
cluckindan · · focus · HN ↗
<a href="https://palisaderesearch.org/research/self-replication" rel="nofollow">https://palisaderesearch.org/research/self-replication