‹ BackHN Continuity

Thread

Understanding the Impact of LLM Watermarking on AI Agent Behavior

58 points · 72 comments · nisosguy

  1. WithinReason · · focus · HN ↗
    This is getting tiring. Watermarking has no effect on model output quality when implemented correctly. It's somewhat like swapping a random RNG seed to the seed 42, and detecting what the seed was from a random sequence. The sequence generated from the seed 42 is just as random as any other seed. There couldn't be a quality difference. And yes, the output from an LLM is a conditional random sequence of tokens from a distribution determined by a model.
    1. lemagedurage · · focus · HN ↗
      That's not true. Watermarks are messing with the next token generation probabilities based on some random seed. The quality is neccesarily lower, the difference is simply too small to notice, typically.
      1. samsartor · · focus · HN ↗
        No, the probability distribution is the same. Watermarking changes the rng sequence used to pick from that distribution.
        1. lemagedurage · · focus · HN ↗
          Theoretically, that can be true.

          In reality, Gemini and Anthropic use SynthID watermarking which affects token probability distribution, i.e. their tournament sampling can pick lower-probability tokens which the LLM's distribution would otherwise not have. They likely use this over unbiased watermarking because SynthID is resistant against text edits.

        2. asnelt · · focus · HN ↗
          The conditional distribution is the same. The joint distribution of the full output does change with watermarking. Just look at the diagram in the article: there is a loop back to the seed once a token has been generated.

          Changing the distribution is the whole point: it introduces statistical regularities that can be detected.

          1. jeremysalwen · · focus · HN ↗
            It introduces statistical regularities, but all RNGs introduce statistical regularities. So the question is on average are these statistical regularities better or worse than those introduced by the alternative, and the answer is no, if implemented properly.
            1. asnelt · · focus · HN ↗
              > It introduces statistical regularities, but all RNGs introduce statistical regularities.

              Agreed.

              > So the question is on average are these statistical regularities better or worse than those introduced by the alternative,

              Indeed. This is an empirical question. That is what the article is about, for a specific setting.

              > ... and the answer is no, if implemented properly.

              I don't agree with that. This is not about implementation. It's about how the statistical regularities that are imposed on the full output distribution affect that distribution. There is a change - by construction. That change can be good in some situations and bad in others. The authors claim that it is mostly bad in the setting that they investigated. This looks like a fair statement.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.