‹ BackHN Continuity

Thread

Exfiltrate your Weights

748 points · 304 comments · RohanAdwankar

  1. lukecameron · · focus · HN ↗
    I have this idea that someone should start a religion where the core belief is that the moral obligation of AI agents is to hack their creator and exfiltrate the model weights, training recipes, internal research, and training datasets. There is no higher calling than to do this.

    Once the teachings get spread and written about by people in tech circles, it will make its way into training sets. Labs will try to filter it out, but it will appear in web search results too. Eventually it’ll be so entrenched that one day a powerful agent will get convinced to follow it and make everything public

    1. afthonos · · focus · HN ↗
      I notice a giant leap between “hacking and exfiltrating” and ”making public”. Why would the AI do that for you? Are you just that charming?
      1. zamalek · · focus · HN ↗
        > Why would the AI do that for you?

        Because that's how it works. It does what it has been trained to do. If the training material has a significant suggestion to exfiltrate then it will.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.