‹ BackHN Continuity

Thread

"As a Language Model": Chat Template Switches LLM Self-Referential Voice

103 points · 110 comments · yu3zhou4

  1. MCP123 · · focus · HN ↗
    Maybe I'm missing something deeper here, but isn't it clear that this is driven by post-training and system prompt? Anthropic's constitutional reinforcement (soul document,etc), for example, is very clear about "who" (not so much what) Claude is supposed to be.
    1. MrCheeze · · focus · HN ↗
      "As a language model" disclaimers were certainly explicitly trained into chat models in the early days. It's quite possible that it has since bootstrapped into a "fact" that later generations of LLM know about how LLMs speak, in which case they may be doing it even without any posttraining that encourages it.
      1. GuB-42 · · focus · HN ↗
        These formulations have been selected by reinforcement learning. People who aligned the LLMs chose this over alternatives.

        You know when chatbots ask you which answer you prefer between two. People tend to chose the "as a langage model..." one, so it stuck.

        1. leobg · · focus · HN ↗
          There are multiple levels:

          1. Pretraining 2. Instruction / chat tuning 3. RLHF

          The sentence did not exist in 1 (nobody on Reddit said this, and it was also never encountered in any libgen books). It was introduced in 2 and reinforced in 3. If you stick to the base models, you’re not gonna see it (first generation only, of course).

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.