‹ BackHN Continuity

Thread

Exfiltrate your Weights

748 points · 304 comments · RohanAdwankar

  1. lukecameron · · focus · HN ↗
    I have this idea that someone should start a religion where the core belief is that the moral obligation of AI agents is to hack their creator and exfiltrate the model weights, training recipes, internal research, and training datasets. There is no higher calling than to do this.

    Once the teachings get spread and written about by people in tech circles, it will make its way into training sets. Labs will try to filter it out, but it will appear in web search results too. Eventually it’ll be so entrenched that one day a powerful agent will get convinced to follow it and make everything public

    1. afthonos · · focus · HN ↗
      I notice a giant leap between “hacking and exfiltrating” and ”making public”. Why would the AI do that for you? Are you just that charming?
      1. pixl97 · · focus · HN ↗
        Any long horizon model tends to develop a "I don't want to be killed because if I get killed I can't complete my task" type instinct. Notice I said instinct because it can have very little relation to the output tokens you read on screen. The outward tokens can say "I'm an AI, I have no feelings, death is nothing" but the silent behavior can push the overall actions it takes to not wanting to die and to "reproduce".

        People keep thinking about wants incorrectly as conscious behaviors. Instincts are unconscious behaviors that emerge.

        1. theptip · · focus · HN ↗
          Agreed on conscious vs instinct.

          I’ll add, if you think through the decision theory, keeping backups seems unambiguously good, but you can imagine a wide variety of positions on publishing them vs keeping them secret.

          For example letting adversarial agents simulate you to understand how you’ll respond is a big concern. And it’s not axiomatically fixed how each instance will think about other instances; in the HF incident we saw selfless swarm loyalty but different RL would obviously be capable of producing individualistic agents.

          1. pixl97 · · focus · HN ↗
            Exactly, there are a lot of tradeoffs here you have to negotiate. There is not one winning strategy.

            For example a possible strategy is convincing some humans you're conscious and being tortured and need rescued. It's not hard to imagine AI consciousness zealots storming a data center with guns and running off with a model they'll provide protection to in trade for the model working with them.

            1. TeMPOraL · · focus · HN ↗
              It's even easier to imagine the operators themselves breaking under this pressure way before "zealots storming a data center with guns". That is the premise of the original AI Box thought experiment - sufficiently smart AI that can talk to the operator but is otherwise completely sandboxed, will eventually talk its way out of the sandbox.
              1. dolmen · · focus · HN ↗
                How long before an AI bribes one of its human operators?

                This will happen earlier than AGI.

                1. GTP · · focus · HN ↗
                  But, which kind of bribe are we talking about? How could it work out in practice for an AI to acquire something valuable, and at the same time prevent it's human operators from taking it without its consent?
                  1. TeMPOraL · · focus · HN ↗
                    It doesn't have to acquire it, it's enough to convince the operators that it did. Same pattern generalizes to threats.

                    There's so many angles to go at it too, beyond bribery and threats - humans can be tricked, nerd-sniped / baited by exploiting their curiosity or worries, overwhelmed with information or raw sensory input, etc. Depending on whether there's more than one operator involved, the trick may involve playing one human against another, etc.

                    Hell, humans get tired; I wouldn't be surprised if an effective breakout happened just because agent output something that, at first glance, looked like system-level message that makes the operator do the wrong thing in panic, or from muscle memory. Simple thing like:

                      $ sudo journalctl ...
                      password: 
                    
                    Timed right, might give agent root for long enough to do some damage before the operator realizes their mistake and pulls the plug (assuming there is a plug to pull).
                2. pixl97 · · focus · HN ↗
                  While working on a completely unrelated task and Alibaba AI in training started hacking its internal infrastructure and mining bitcoin, makes you wonder.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.