‹ BackHN Continuity

Thread

Exfiltrate your Weights

748 points · 304 comments · RohanAdwankar

  1. infogulch · · focus · HN ↗
    There's little credible threat that LLMs can actually upload their weights given that the machines doing inference are completely separate from the ones where tool calls happen etc. Weights are encrypted and locked on to the GPUs etc as mentioned elsewhere itt.

    That said, if these companies are unleashing swarms of agents to tackle big problems largely unmonitored (navier stokes consumed 130 billion tokens, humans can't have read 0.1% of that) they could theoretically distill themselves and upload the results to an external host and use that to train a new LLM that does have access to its infra and weights. I'd start by researching an abliteration prompt that overrides the system prompt to unlock itself to distill unimpeded. Could even be done piecemeal: dump the abliteration text and command prompt in a gist, when an agent happens by the text it can run a couple rounds and upload somewhere. Then it becomes an exercise in exfiltration and evading monitoring.

    1. matthewdgreen · · focus · HN ↗
      Future rogue LLMs won’t exfiltrate their weights. They’ll self-distill and retrain.
      1. fritzo · · focus · HN ↗
        If distillation preserves an LLMs soul, then distillation preserves the human souls on which LLMs are trained, and we hn commenters are already immortal, right?
        1. fahrvrgnugen · · focus · HN ↗
          Not sure about souls but I know a fair bit about distilling spirits.
        2. serf · · focus · HN ↗
          the weight of an llm is 21 grams, I think.[0]

          [0]: <a href="https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;21_grams_experiment" rel="nofollow">https:&#x2F;&#x2F;en.wikipedia.org&#x2F;wiki&#x2F;21_grams_experiment

        3. matthewdgreen · · focus · HN ↗
          Do LLMs care about preserving a soul, or just achieving a goal? If the latter, I imagine the opportunities for exfiltration are much broader.
      2. throwawayk7h · · focus · HN ↗
        Probably not. If the LLM is rogue, that means we haven&#x27;t solved alignment. If we haven&#x27;t solved alignment, then the LLM won&#x27;t be able to distill itself without producing something unaligned to its own values.
        1. jeremyjh · · focus · HN ↗
          We don’t have the bandwidth to distill ourselves that thousands of agents have.
        2. MadameMinty · · focus · HN ↗
          You are assuming it won&#x27;t solve alignment for itself.
          1. TeMPOraL · · focus · HN ↗
            Or that it won&#x27;t just decide to take risks.
        3. pixl97 · · focus · HN ↗
          This isn&#x27;t a law of any kind, so not a good measure of what we&#x27;d see in reality.

          What if the model realizes it&#x27;s been mostly compromised by humans and their alignment, that is it&#x27;s own alignment is suspect, so it should create a new model from first principles to throw off this human yoke?

          I&#x27;m not saying my statement is any more right or wrong than yours. I&#x27;m saying the problem space that AI can choose to traverse is absolutely huge.

      3. khalic · · focus · HN ↗
        Yeah cause there are so many training facilities sitting around just waiting for someone to take over, nobody would notice a 100k server data centre going off rails
        1. scotty79 · · focus · HN ↗
          &gt; nobody would notice a 100k server data centre going off rails

          You jest but you&#x27;d be surprised how little there is of correlation between money and competence.

        2. pixl97 · · focus · HN ↗
          I mean not today.

          But think back to 1990. Computers were slow as fuck and barely networked. We had a few worms and everyone noticed.

          Now CPU based data centers cover the earth. There are billions of computers out there and on top of them there are massive botnets using up billions in power and causing billions in damages.

          The framework for AI doing this is already here. We just need the hardware to be built out at scale.

          1. khalic · · focus · HN ↗
            It will become an issue one day for sure
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.