The Implications of Linguistic Illegibility for LLM Security
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
The Implications of Linguistic Illegibility for LLM Security
Unofficial Hacker News client; not affiliated with Y Combinator.
mnkv · · focus · HN ↗
I dislike this term because it doesn't explain where this "illegibility" is coming from. Models are post-trained towards non-linguistic goals with (mostly) non-linguistic rewards. A model's reasoning chain is reinforced if it leads to a correct answer or agentic goal. It doesn't need to be linguistically accurate and meanings can drift over training.
trhway · · focus · HN ↗