"As a Language Model": Chat Template Switches LLM Self-Referential Voice
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
"As a Language Model": Chat Template Switches LLM Self-Referential Voice
Unofficial Hacker News client; not affiliated with Y Combinator.
LiamPowell · · focus · HN ↗
Presumably the fact that they're heavily trained to reply in this way? I don't know about the rest of the paper, but this part sticks out as a really odd claim unless I'm entirely misunderstanding this part.
yu3zhou4 · · focus · HN ↗
bonoboTP · · focus · HN ↗
Those don't have a themselves, because they can only continue text. A base model can only plausibly continue along the lines of what a character would say in a novel or what the narration would say in a story or in an article. Post-trained models may tie "I"-talk to actually observable effects they caused in some RL environment, or to how RLHF humans rewards its self-talk. But there is no themselves in a base model.