‹ BackHN Continuity

Thread

"As a Language Model": Chat Template Switches LLM Self-Referential Voice

103 points · 110 comments · yu3zhou4

  1. Izmaki · · focus · HN ↗
    "As a Language Model..." is one of the beginnings of a sentence I hate the most from LLMs and is the reason why I support free (as in "Liberty"), local models. I'm well aware that it is not a doctor and cannot replace a real doctor with multiple years of experience, I don't need to waste braincell activity on reading that it "as a Language Model" cannot give a precise diagnosis and that I should ask a real doctor - all I want to know is if I what I experience justifies either A) ER, B) 3-4 weeks scheduled doctors appointment or C) two paracetamol and a nap.

    I don't want "jailbroken" LLMs to commit crime. I want them to avoid having this vendor-specific "bloatware" all over the product I'm using.

    1. Anduia · · focus · HN ↗
      Be careful there. LLMs may be good at identifying a condition based on a description of the symptoms, but they are much worse at recommending the correct course of action (getting it wrong half of the time).

      [0] <a href="https:&#x2F;&#x2F;www.nature.com&#x2F;articles&#x2F;s41591-025-04074-y" rel="nofollow">https:&#x2F;&#x2F;www.nature.com&#x2F;articles&#x2F;s41591-025-04074-y

      1. Izmaki · · focus · HN ↗
        ...I know, which is why &quot;as a Language Model and not a real doctor&quot; is a pointless comment to start off with. It should simply not recommend treatment if it&#x27;s not sure it is correct. I wouldn&#x27;t blame it or anyone if they asked for help treating a stiff neck, and the LLM (or your neighbor or parent or spouse) suggested light exercises to help relieve it - and do not jump to the suspicion that you may have meningitis.

        As a Human, I do not need to know it is a Language Model.

        1. StilesCrisis · · focus · HN ↗
          LLMs are famously bad at determining &quot;if it&#x27;s not sure it is correct.&quot; They are always confident, because a confident tone ranks better in RL.
          1. wxnx · · focus · HN ↗
            &gt; They are always confident, because a confident tone ranks better in RL.

            This makes it sound like RL rewards a confident tone -- in general, I don&#x27;t think this is true (most RL is RLVR, which typically uses binary verification of correctness).

            I say this because the real reason &quot;they are always confident&quot; is in some sense even more contrived. Training text where the speaker sounded more confident is more likely to contain a correct answer.

            1. Forgeties79 · · focus · HN ↗
              &gt; This makes it sound like RL rewards a confident tone

              Generally it does. Especially in groups. Hell look at the state of politics right now: it’s basically about being the loudest, least compromising, most confident voice in the room. It’s not just because people will assume you’re correct, it’s because if you are confidently saying something that someone wants to be right, then they’re often just going to follow it. We are all guilty of this.

              If I’m turning to an LLM to diagnose something medical, I am probably frustrated or uncomfortable. Maybe I’m just scared. So this magic device just instantly spits out (allegedly) exactly what is wrong and exactly what I need to do with no hesitation. I am very liable to just take it at face value because I want an answer and it gave me one, as we have seen over and over again since ChatGPT was unleashed on the world.

              We don’t really need to speculate, this is already a problem.

              1. wxnx · · focus · HN ↗
                Sorry, I understand now we&#x27;re talking about different things. You&#x27;re talking about preference optimization.

                I was unintentionally being pedantic, because this isn&#x27;t really done with RL anymore - it doesn&#x27;t need to be. RL is now typically only used to train reasoning for tasks with a well-defined correct answer (that&#x27;s what I meant by binary reward) - this is called RLVR (RL with verifiable rewards).

                Preference optimization (training the model on user &quot;this response is better than that response&quot; type data) is more often done with something in the same family as DPO (direct preference optimization), which is decidedly not RL.

                Your philosophical concerns are correct of course. And there&#x27;s the added caveat that the models that most people are using are closed, so we don&#x27;t actually know their training recipes for sure.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.