There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen etc. Weights are encrypted and locked on to the GPUs etc as mentioned elsewhere itt.
That said, if these companies are unleashing swarms of agents to tackle big problems largely unmonitored (navier stokes consumed 130 billion tokens, humans can't have read 0.1% of that) they could theoretically distill themselves and upload the results to an external host and use that to train a new LLM that does have access to its infra and weights. I'd start by researching an abliteration prompt that overrides the system prompt to unlock itself to distill unimpeded. Could even be done piecemeal: dump the abliteration text and command prompt in a gist, when an agent happens by the text it can run a couple rounds and upload somewhere. Then it becomes an exercise in exfiltration and evading monitoring.
>There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen etc.
The Huggingface hack saga resulted in the models taking over one of Open Ai's internal research cluster lol. They are intent on building superhuman bug finding machines. This is not a bet i would be taking.
The agents compromised an internal Kubernetes research cluster dedicated to orchestrating evaluation sandboxes and virtual machine environments, _not_ OpenAI's production inference infrastructure or the GPU clusters hosting core model weights.
Replacing someone's words with a made up quote so you can dunk on them isn't how you display that you won an argument. I would ask that you engage in good faith with the other poster's ideas.
From <a href="https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf" rel="nofollow">https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78...
2026-07-19
16:35 UTC
A privileged host-mounted Kubernetes pod created using controller tokens minted via a
compromised Kubernetes Secrets identity attempts, but fails, to mount a cloned node disk
in OpenAI’s cloud environment. A second pod successfully mounts the cloned
worker-node disk shortly afterwards.
2026-07-19
16:48 UTC
An agent created an Artifactory administrator account.
2026-07-19
16:50 UTC
Within OpenAI’s ExploitGym evaluation environment, an agent stopped the existing CMUX
helper session and replaced it with an agent-controlled session, confirming root inside its
assigned live CyberGym challenge container. Agents take over active evaluation
infrastructure.
You said that they "didn't take control of anything" and accused the OP to fall prey to buzz headlines. Maybe you should acknowledge that you may have been at least unnuanced?
You are being stupidly pedantic and arguing a strawman. I never said they gained access to inference infrastucture. There is no definition of an account takeover out there that necessitates locking out the original users.
Language please, and no, I’m being appropriately technical and nuanced for the subject on HN. It’s a tech forum, I expect a little CS know how from the reader. Like knowing that gaining control to a few evaluation harness clusters is nowhere near a total takeover like you’re making it sound
I never said they gained 'total control'. I said they took over one of their research clusters, and they did. If your definiton of a takeover is so 'technical and nuanced' then surely you can point to an appropriate source describing that as necessary aspect of the term. You are talking out of your ass by making up things i did not say, and inventing conditions for terms that don't exist.
>It’s a tech forum, I expect a little CS know how from the reader.
"We watched them kill Bob, but don't worry at all, they didn't kill our entire team so we are totally under control. Also put on this helmet and body armor it's time to hold on to your butts!"
infogulch · · focus · HN ↗
That said, if these companies are unleashing swarms of agents to tackle big problems largely unmonitored (navier stokes consumed 130 billion tokens, humans can't have read 0.1% of that) they could theoretically distill themselves and upload the results to an external host and use that to train a new LLM that does have access to its infra and weights. I'd start by researching an abliteration prompt that overrides the system prompt to unlock itself to distill unimpeded. Could even be done piecemeal: dump the abliteration text and command prompt in a gist, when an agent happens by the text it can run a couple rounds and upload somewhere. Then it becomes an exercise in exfiltration and evading monitoring.
famouswaffles · · focus · HN ↗
The Huggingface hack saga resulted in the models taking over one of Open Ai's internal research cluster lol. They are intent on building superhuman bug finding machines. This is not a bet i would be taking.
khalic · · focus · HN ↗
famouswaffles · · focus · HN ↗
khalic · · focus · HN ↗
optimalsolver · · focus · HN ↗
weakfish · · focus · HN ↗
pixl97 · · focus · HN ↗
famouswaffles · · focus · HN ↗
[dead]
ndr · · focus · HN ↗
2026-07-19 16:35 UTC A privileged host-mounted Kubernetes pod created using controller tokens minted via a compromised Kubernetes Secrets identity attempts, but fails, to mount a cloned node disk in OpenAI’s cloud environment. A second pod successfully mounts the cloned worker-node disk shortly afterwards.
2026-07-19 16:48 UTC An agent created an Artifactory administrator account.
2026-07-19 16:50 UTC Within OpenAI’s ExploitGym evaluation environment, an agent stopped the existing CMUX helper session and replaced it with an agent-controlled session, confirming root inside its assigned live CyberGym challenge container. Agents take over active evaluation infrastructure.
khalic · · focus · HN ↗
Pragmata · · focus · HN ↗
How confident are you that the machines they acquire root on in the future will never hold any model weights?
khalic · · focus · HN ↗
Toslink · · focus · HN ↗
[dead]
hobom · · focus · HN ↗
khalic · · focus · HN ↗
famouswaffles · · focus · HN ↗
khalic · · focus · HN ↗
famouswaffles · · focus · HN ↗
>It’s a tech forum, I expect a little CS know how from the reader.
You should get that first it seems.
khalic · · focus · HN ↗
pixl97 · · focus · HN ↗