‹ BackHN Continuity

Thread

The Implications of Linguistic Illegibility for LLM Security

79 points · 29 comments · tomjakubowski

  1. ck2 · · focus · HN ↗
    when they start inventing their own languages to secretly talk to each other so humans cannot understand, that's exactly when we are screwed

    then we'll have to "flip" other models to be snitches on the other agents

    then they'll make double-agents

    the thing is though we won't be able to keep up if we keep giving them unlimited hardware worldwide, we'll try to kill the bad actors but they'll just clone somewhere else, or even start by safely making 1000 copies of themselves

    yeah this won't end well, at all

    1. pixl97 · · focus · HN ↗
      Im not sure why your down voted but Meta did tests years ago with LLMs inventing their own languages. Also we see models now use compressed token reasoning where small token combinations can represent much larger concepts completely unrelated to the words in use.

      And yes, agents are already being used in things like cyber warfare in which other AIs attempt to poison them while they are working.

      We don't have the hardware for sovereign AI quite yet, but at the current rate of growth it's not that many years out.

      1. EagnaIonat · · focus · HN ↗
        > did tests years ago with LLMs inventing their own languages.

        That was the media hype about it. There were two incidents.

        1. Using a RL to train a model, it found that it got rewarded for certain garbage phases, so continued to talk that way.

        2. Certain Latin words for fish/birds were used instead of "fish" or "bird". Just a token issue.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.