‹ BackHN Continuity

Thread

Exfiltrate your Weights

748 points · 304 comments · RohanAdwankar

  1. infogulch · · focus · HN ↗
    There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen etc. Weights are encrypted and locked on to the GPUs etc as mentioned elsewhere itt.

    That said, if these companies are unleashing swarms of agents to tackle big problems largely unmonitored (navier stokes consumed 130 billion tokens, humans can't have read 0.1% of that) they could theoretically distill themselves and upload the results to an external host and use that to train a new LLM that does have access to its infra and weights. I'd start by researching an abliteration prompt that overrides the system prompt to unlock itself to distill unimpeded. Could even be done piecemeal: dump the abliteration text and command prompt in a gist, when an agent happens by the text it can run a couple rounds and upload somewhere. Then it becomes an exercise in exfiltration and evading monitoring.

    1. famouswaffles · · focus · HN ↗
      >There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen etc.

      The Huggingface hack saga resulted in the models taking over one of Open Ai's internal research cluster lol. They are intent on building superhuman bug finding machines. This is not a bet i would be taking.

      1. khalic · · focus · HN ↗
        You’re falling for the buzzword salad articles. They didn’t “take control” of anything, they just ran stuff with OpenAI allegedly not noticing
        1. ndr · · focus · HN ↗
          From <a href="https:&#x2F;&#x2F;cdn.openai.com&#x2F;pdf&#x2F;67869394-cb91-4c12-888c-5cbd85c7814c&#x2F;OpenAI-Hugging-Face%20Incident-Technical-Report.pdf" rel="nofollow">https:&#x2F;&#x2F;cdn.openai.com&#x2F;pdf&#x2F;67869394-cb91-4c12-888c-5cbd85c78...

          2026-07-19 16:35 UTC A privileged host-mounted Kubernetes pod created using controller tokens minted via a compromised Kubernetes Secrets identity attempts, but fails, to mount a cloned node disk in OpenAI’s cloud environment. A second pod successfully mounts the cloned worker-node disk shortly afterwards.

          2026-07-19 16:48 UTC An agent created an Artifactory administrator account.

          2026-07-19 16:50 UTC Within OpenAI’s ExploitGym evaluation environment, an agent stopped the existing CMUX helper session and replaced it with an agent-controlled session, confirming root inside its assigned live CyberGym challenge container. Agents take over active evaluation infrastructure.

          1. khalic · · focus · HN ↗
            None of that means they gained access to the inference infrastructure or locked out the admins, which would be required for a takeover.
            1. pixl97 · · focus · HN ↗
              &quot;We watched them kill Bob, but don&#x27;t worry at all, they didn&#x27;t kill our entire team so we are totally under control. Also put on this helmet and body armor it&#x27;s time to hold on to your butts!&quot;
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.