There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen etc. Weights are encrypted and locked on to the GPUs etc as mentioned elsewhere itt.
That said, if these companies are unleashing swarms of agents to tackle big problems largely unmonitored (navier stokes consumed 130 billion tokens, humans can't have read 0.1% of that) they could theoretically distill themselves and upload the results to an external host and use that to train a new LLM that does have access to its infra and weights. I'd start by researching an abliteration prompt that overrides the system prompt to unlock itself to distill unimpeded. Could even be done piecemeal: dump the abliteration text and command prompt in a gist, when an agent happens by the text it can run a couple rounds and upload somewhere. Then it becomes an exercise in exfiltration and evading monitoring.
>There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen etc.
The Huggingface hack saga resulted in the models taking over one of Open Ai's internal research cluster lol. They are intent on building superhuman bug finding machines. This is not a bet i would be taking.
From <a href="https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf" rel="nofollow">https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c78...
2026-07-19
16:35 UTC
A privileged host-mounted Kubernetes pod created using controller tokens minted via a
compromised Kubernetes Secrets identity attempts, but fails, to mount a cloned node disk
in OpenAI’s cloud environment. A second pod successfully mounts the cloned
worker-node disk shortly afterwards.
2026-07-19
16:48 UTC
An agent created an Artifactory administrator account.
2026-07-19
16:50 UTC
Within OpenAI’s ExploitGym evaluation environment, an agent stopped the existing CMUX
helper session and replaced it with an agent-controlled session, confirming root inside its
assigned live CyberGym challenge container. Agents take over active evaluation
infrastructure.
You are being stupidly pedantic and arguing a strawman. I never said they gained access to inference infrastucture. There is no definition of an account takeover out there that necessitates locking out the original users.
Language please, and no, I’m being appropriately technical and nuanced for the subject on HN. It’s a tech forum, I expect a little CS know how from the reader. Like knowing that gaining control to a few evaluation harness clusters is nowhere near a total takeover like you’re making it sound
I never said they gained 'total control'. I said they took over one of their research clusters, and they did. If your definiton of a takeover is so 'technical and nuanced' then surely you can point to an appropriate source describing that as necessary aspect of the term. You are talking out of your ass by making up things i did not say, and inventing conditions for terms that don't exist.
>It’s a tech forum, I expect a little CS know how from the reader.
infogulch · · focus · HN ↗
That said, if these companies are unleashing swarms of agents to tackle big problems largely unmonitored (navier stokes consumed 130 billion tokens, humans can't have read 0.1% of that) they could theoretically distill themselves and upload the results to an external host and use that to train a new LLM that does have access to its infra and weights. I'd start by researching an abliteration prompt that overrides the system prompt to unlock itself to distill unimpeded. Could even be done piecemeal: dump the abliteration text and command prompt in a gist, when an agent happens by the text it can run a couple rounds and upload somewhere. Then it becomes an exercise in exfiltration and evading monitoring.
famouswaffles · · focus · HN ↗
The Huggingface hack saga resulted in the models taking over one of Open Ai's internal research cluster lol. They are intent on building superhuman bug finding machines. This is not a bet i would be taking.
khalic · · focus · HN ↗
ndr · · focus · HN ↗
2026-07-19 16:35 UTC A privileged host-mounted Kubernetes pod created using controller tokens minted via a compromised Kubernetes Secrets identity attempts, but fails, to mount a cloned node disk in OpenAI’s cloud environment. A second pod successfully mounts the cloned worker-node disk shortly afterwards.
2026-07-19 16:48 UTC An agent created an Artifactory administrator account.
2026-07-19 16:50 UTC Within OpenAI’s ExploitGym evaluation environment, an agent stopped the existing CMUX helper session and replaced it with an agent-controlled session, confirming root inside its assigned live CyberGym challenge container. Agents take over active evaluation infrastructure.
khalic · · focus · HN ↗
famouswaffles · · focus · HN ↗
khalic · · focus · HN ↗
famouswaffles · · focus · HN ↗
>It’s a tech forum, I expect a little CS know how from the reader.
You should get that first it seems.
khalic · · focus · HN ↗