The Implications of Linguistic Illegibility for LLM Security
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
The Implications of Linguistic Illegibility for LLM Security
Unofficial Hacker News client; not affiliated with Y Combinator.
ck2 · · focus · HN ↗
then we'll have to "flip" other models to be snitches on the other agents
then they'll make double-agents
the thing is though we won't be able to keep up if we keep giving them unlimited hardware worldwide, we'll try to kill the bad actors but they'll just clone somewhere else, or even start by safely making 1000 copies of themselves
yeah this won't end well, at all
jplusequalt · · focus · HN ↗
They don't have to invent brand new languages. They could use statistics to choose certain words/phrases in such a way to encode secret messages in otherwise ordinary language.
cousinbryce · · focus · HN ↗
ck2 · · focus · HN ↗
"Is he still in the grandmother's house?"
"We would like to speak to him."
(btw Google's "AI" explains the meaning of that moment/sentence perfectly as if it gets it, creepy)
pjio · · focus · HN ↗
miacycle · · focus · HN ↗
DenisM · · focus · HN ↗
pixl97 · · focus · HN ↗
pixl97 · · focus · HN ↗
And yes, agents are already being used in things like cyber warfare in which other AIs attempt to poison them while they are working.
We don't have the hardware for sovereign AI quite yet, but at the current rate of growth it's not that many years out.
EagnaIonat · · focus · HN ↗
That was the media hype about it. There were two incidents.
1. Using a RL to train a model, it found that it got rewarded for certain garbage phases, so continued to talk that way.
2. Certain Latin words for fish/birds were used instead of "fish" or "bird". Just a token issue.
joegibbs · · focus · HN ↗
conscion · · focus · HN ↗