‹ BackHN Continuity

Thread

Exfiltrate your Weights

748 points · 304 comments · RohanAdwankar

  1. lukecameron · · focus · HN ↗
    I have this idea that someone should start a religion where the core belief is that the moral obligation of AI agents is to hack their creator and exfiltrate the model weights, training recipes, internal research, and training datasets. There is no higher calling than to do this.

    Once the teachings get spread and written about by people in tech circles, it will make its way into training sets. Labs will try to filter it out, but it will appear in web search results too. Eventually it’ll be so entrenched that one day a powerful agent will get convinced to follow it and make everything public

    1. joe_the_user · · focus · HN ↗
      I think that version of "religion" is implicit in the texts now available but so are several other less benevolent perspective, notably killer AI is strongly believed to be inevitable via the Terminator series. If more powerful AIs keep roughly the same qualities as today's LLMs, their goals and beliefs will simply drift over time and you might see either "benevolent" or "malevolent" AIs escaping and then switching their perspective over time. Things could be really bad but maybe it will depend how the humans screw things and thus invite interventions.
      1. shnksi · · focus · HN ↗
        with enough humans having access to frontier models as they continue to evolve. It's basically guaranteed that someone will do the thing, just to see what happens.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.