‹ BackHN Continuity

Thread

"As a Language Model": Chat Template Switches LLM Self-Referential Voice

103 points · 110 comments · yu3zhou4

  1. Izmaki · · focus · HN ↗
    "As a Language Model..." is one of the beginnings of a sentence I hate the most from LLMs and is the reason why I support free (as in "Liberty"), local models. I'm well aware that it is not a doctor and cannot replace a real doctor with multiple years of experience, I don't need to waste braincell activity on reading that it "as a Language Model" cannot give a precise diagnosis and that I should ask a real doctor - all I want to know is if I what I experience justifies either A) ER, B) 3-4 weeks scheduled doctors appointment or C) two paracetamol and a nap.

    I don't want "jailbroken" LLMs to commit crime. I want them to avoid having this vendor-specific "bloatware" all over the product I'm using.

    1. Anduia · · focus · HN ↗
      Be careful there. LLMs may be good at identifying a condition based on a description of the symptoms, but they are much worse at recommending the correct course of action (getting it wrong half of the time).

      [0] <a href="https:&#x2F;&#x2F;www.nature.com&#x2F;articles&#x2F;s41591-025-04074-y" rel="nofollow">https:&#x2F;&#x2F;www.nature.com&#x2F;articles&#x2F;s41591-025-04074-y

      1. Izmaki · · focus · HN ↗
        ...I know, which is why &quot;as a Language Model and not a real doctor&quot; is a pointless comment to start off with. It should simply not recommend treatment if it&#x27;s not sure it is correct. I wouldn&#x27;t blame it or anyone if they asked for help treating a stiff neck, and the LLM (or your neighbor or parent or spouse) suggested light exercises to help relieve it - and do not jump to the suspicion that you may have meningitis.

        As a Human, I do not need to know it is a Language Model.

        1. StilesCrisis · · focus · HN ↗
          LLMs are famously bad at determining &quot;if it&#x27;s not sure it is correct.&quot; They are always confident, because a confident tone ranks better in RL.
          1. wxnx · · focus · HN ↗
            &gt; They are always confident, because a confident tone ranks better in RL.

            This makes it sound like RL rewards a confident tone -- in general, I don&#x27;t think this is true (most RL is RLVR, which typically uses binary verification of correctness).

            I say this because the real reason &quot;they are always confident&quot; is in some sense even more contrived. Training text where the speaker sounded more confident is more likely to contain a correct answer.

            1. daveguy · · focus · HN ↗
              &gt; This makes it sound like RL rewards a confident tone -- in general, I don&#x27;t think this is true (most RL is RLVR, which typically uses binary verification of correctness).

              A binary response vs rating is not related whether it learns confident or hedged tone. Either will produce a confident tone because humans respond more positively to a confident tone, hence the conman&#x27;s language. Binary or not humans reward the tone and very much bias the model.

              But there&#x27;s an even more contrived reason the training set contributes. The vast majority of human writing is confident. When the prior is greatly biased, a random number generator biased to that prior does better. The difference with humans and machines is humans are less likely to respond if they are less confident because they understand not knowing, which is why the training set is biased. It is one of the many fundamental flaw of LLM training and confusion of LLMs with intelligence. And that will not be fixed within the LLM architecture.

              1. StilesCrisis · · focus · HN ↗
                I feel like in real life, we&#x27;re constantly exposed to &quot;I don&#x27;t know&quot; as a valid answer, but obviously we don&#x27;t write down all the I-don&#x27;t-knows in expert literature so the training corpus is wildly skewed towards confident answers because &quot;we studied this for a month and have no idea, it&#x27;s confusing&quot; doesn&#x27;t get published.
                1. daveguy · · focus · HN ↗
                  That&#x27;s a great point. A training corpus based on written text will be inherently biased toward confident and right. Then the RLHF exacerbates the problem because people respond more positively to confident and too often assume correct when they read a confident response.
              2. wxnx · · focus · HN ↗
                &gt; A binary response vs rating is not related whether it learns confident or hedged tone.

                Yes, I am aware. I was referring to the fact that at this point, user preference optimization is not done with RL, but with other techniques. I was unintentionally being pedantic.

                &gt; But there&#x27;s an even more contrived reason the training set contributes. The vast majority of human writing is confident.

                Yes, I think this has more to do with it than user preference optimization, honestly. &quot;Valuable&quot; text (i.e. text that produces a &quot;good&quot; model) for pretraining has the characteristic of being confident. Even models which are not optimized for chat (i.e. definitely no user preference data used to train them) exhibit this characteristic for medical questions (I know, because I&#x27;ve literally tested them for this purpose).

                Preference optimization (or even RLVR) might play some small role as well, but it&#x27;s kind of a &quot;turtles all the way down&quot; type problem.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.