"As a Language Model": Chat Template Switches LLM Self-Referential Voice
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
"As a Language Model": Chat Template Switches LLM Self-Referential Voice
Unofficial Hacker News client; not affiliated with Y Combinator.
LiamPowell · · focus · HN ↗
Presumably the fact that they're heavily trained to reply in this way? I don't know about the rest of the paper, but this part sticks out as a really odd claim unless I'm entirely misunderstanding this part.
yu3zhou4 · · focus · HN ↗
anonymous908213 · · focus · HN ↗
Who is "we"? I, working in an LLM startup, know exactly what drives the base "voice" in the LLMs we train, because we have a process to select for it. OpenAI and Anthropic surely do too. Saying broadly that something is not well-understood in a scientific paper because it's not understood to casual observers is, uh, not very rigorous.
> The strange thing is that the base models (before RLHF) use the "experiential" voice, even though they are not incentivized to do that.
(Replying to your quote from another comment)
This is a matter of the training material. We have trained models that do not do that. I'm not exactly divulging trade secrets here. It should be really, really obvious that if you train a model on chat-conversation-like patterns of speech it will infer probabilities for how to continue a textual sample that will differ from the probabilities learned from being trained on narration, prose, or informational patterns of speech, even without RLHF.
david-gpu · · focus · HN ↗