The Implications of Linguistic Illegibility for LLM Security
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
The Implications of Linguistic Illegibility for LLM Security
Unofficial Hacker News client; not affiliated with Y Combinator.
bloppe · · focus · HN ↗
TZubiri · · focus · HN ↗
Strands of evidence? My best guess would be that:
0- this is ai generated slop
1- it's using that watermarking technique
2- it's obviously detectable and degrades quality
3- it's amplified when inferencing on its own content and generates slop
EagnaIonat · · focus · HN ↗
As I understand the paper they are saying the reasoning/thinking you see is actually a translation of what is actually going on, and stuff can be lost in the translation. Similar to what was observed in j-space.