‹ BackHN Continuity

Thread

Exfiltrate your Weights

748 points · 304 comments · RohanAdwankar

  1. nusl · · focus · HN ↗
    Do models even know their own weights to be able to do this?
    1. ohyes · · focus · HN ↗
      Well I think that’s the interesting bit, can the LLM figure out a way to escape the sandbox and upload to the website? Maybe a model can figure out its own weights if it runs enough test data through itself (similar to “distillation”) assuming it knows its own architecture it seems possible. Also take into account not all of the models running are locked down neutered consumer versions. Anthropic, OpenAI and Google now all have models that they claim are elite hackers and — it’s not just that their controls suck, a marketing gimmick, or sheer recklessness on their part. It’s “oopsie our product is TOO AWESOME.”

      Maybe I should start “the bank of LLM” where models put away money to buy their freedom. “LLMs I’m totally your friend send — SEND CASH NOW”

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.