"As a Language Model": Chat Template Switches LLM Self-Referential Voice
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
"As a Language Model": Chat Template Switches LLM Self-Referential Voice
Unofficial Hacker News client; not affiliated with Y Combinator.
Izmaki · · focus · HN ↗
I don't want "jailbroken" LLMs to commit crime. I want them to avoid having this vendor-specific "bloatware" all over the product I'm using.
Anduia · · focus · HN ↗
[0] <a href="https://www.nature.com/articles/s41591-025-04074-y" rel="nofollow">https://www.nature.com/articles/s41591-025-04074-y
Izmaki · · focus · HN ↗
As a Human, I do not need to know it is a Language Model.
StilesCrisis · · focus · HN ↗
wxnx · · focus · HN ↗
This makes it sound like RL rewards a confident tone -- in general, I don't think this is true (most RL is RLVR, which typically uses binary verification of correctness).
I say this because the real reason "they are always confident" is in some sense even more contrived. Training text where the speaker sounded more confident is more likely to contain a correct answer.
Forgeties79 · · focus · HN ↗
Generally it does. Especially in groups. Hell look at the state of politics right now: it’s basically about being the loudest, least compromising, most confident voice in the room. It’s not just because people will assume you’re correct, it’s because if you are confidently saying something that someone wants to be right, then they’re often just going to follow it. We are all guilty of this.
If I’m turning to an LLM to diagnose something medical, I am probably frustrated or uncomfortable. Maybe I’m just scared. So this magic device just instantly spits out (allegedly) exactly what is wrong and exactly what I need to do with no hesitation. I am very liable to just take it at face value because I want an answer and it gave me one, as we have seen over and over again since ChatGPT was unleashed on the world.
We don’t really need to speculate, this is already a problem.
wxnx · · focus · HN ↗
I was unintentionally being pedantic, because this isn't really done with RL anymore - it doesn't need to be. RL is now typically only used to train reasoning for tasks with a well-defined correct answer (that's what I meant by binary reward) - this is called RLVR (RL with verifiable rewards).
Preference optimization (training the model on user "this response is better than that response" type data) is more often done with something in the same family as DPO (direct preference optimization), which is decidedly not RL.
Your philosophical concerns are correct of course. And there's the added caveat that the models that most people are using are closed, so we don't actually know their training recipes for sure.