‹ BackHN Continuity

Thread

Understanding the Impact of LLM Watermarking on AI Agent Behavior

58 points · 72 comments · nisosguy

  1. WithinReason · · focus · HN ↗
    This is getting tiring. Watermarking has no effect on model output quality when implemented correctly. It's somewhat like swapping a random RNG seed to the seed 42, and detecting what the seed was from a random sequence. The sequence generated from the seed 42 is just as random as any other seed. There couldn't be a quality difference. And yes, the output from an LLM is a conditional random sequence of tokens from a distribution determined by a model.
    1. jannyfer · · focus · HN ↗
      “When implemented correctly” is probably what people are complaining about.

      Opus 5 started adding a bunch of comments to code, even when instructed not to, and for very simple changes where the comment itself was longer than the code change. Was that so that there are enough tokens outputted for watermarking? Many people suspected so.

      1. inopinatus · · focus · HN ↗
        I have seen Fable’s reasoning talk itself into ignoring an unequivocal prompt directive not to write comments, and then be startled by the precommit hook that rejects it. It is almost desperately predisposed to emit prose. And horribly turgid, waffling prose, to boot. Claude has been like this since Opus 4.7 though, i.e. (probably) predating the introduction of watermarking.
        1. datadrivenangel · · focus · HN ↗
          yeah Anthropic lost the sauce with Opus 4.7 and beyond.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.