‹ BackHN Continuity

Thread

Understanding the Impact of LLM Watermarking on AI Agent Behavior

58 points · 72 comments · nisosguy

  1. WithinReason · · focus · HN ↗
    This is getting tiring. Watermarking has no effect on model output quality when implemented correctly. It's somewhat like swapping a random RNG seed to the seed 42, and detecting what the seed was from a random sequence. The sequence generated from the seed 42 is just as random as any other seed. There couldn't be a quality difference. And yes, the output from an LLM is a conditional random sequence of tokens from a distribution determined by a model.
    1. arcticbull · · focus · HN ↗
      Model companies are doing this for themselves anyways, it’s so they don’t feed generated content back into the slopper and collapse the model. From that angle it over time contributes to better model quality.
      1. Phemist · · focus · HN ↗
        Also - it forms a cartel.

        Detection of watermarking requires access to the watermarking key, a secret in the current suggested scheme (leaking it would amount to being able to strip the watermark).

        So, there will need to be a watermark checking service. The checking service will of course be rate-limited for common folk (and model distillers). OpenAI/Anthropic/Google/other privileged model builders need to filter out AI slop at scale, so need access to others' service without rate-limits (or the watermarking keys need to be shared).

        This creates an in-group with pristine datasets, and an outgroup whose models will collapse on the slop outputs with no good ability to filter.

        1. gruez · · focus · HN ↗
          >This creates an in-group with pristine datasets, and an outgroup whose models will collapse on the slop outputs with no good ability to filter.

          But all the chinese labs who are hot on the heels of american labs thanks to "distillation" seems to be able to work without "pristine datasets"?

          1. Phemist · · focus · HN ↗
            Haha, calling them pristine is maybe too much indeed. Right now they seem to manage, but what % of the scrapable web is now AI slop? What if it becomes 99% AI slop, 99.9%, 99.99%, etc. Surely the signal to slop at some point starts to become too low and you need to do some kind of at-scale filtering.
      2. porridgeraisin · · focus · HN ↗
        What? It has absolutely nothing to do with "model collapse".
        1. WD-42 · · focus · HN ↗
          Do you have any evidence that future training would ignore the presence of a watermark? Seems like a pretty valuable signal to me.
          1. porridgeraisin · · focus · HN ↗
            I really don't see why it matters
            1. WD-42 · · focus · HN ↗
              If the goal of training is to improve weights, training on output of the existing weights won’t improve anything, in fact the opposite may happen.
              1. charcircuit · · focus · HN ↗
                >training on output of the existing weights won’t improve anything

                This is simply false. You are underestimating the utility of synthetic data and the ability to learn from the mistakes the current weights make.

                1. WD-42 · · focus · HN ↗
                  Synthetic data is used in specific contexts. Slopped up hacker news comments along side natural ones is where watermarking will be used to delimitate them. Not all synthetic data is good.
      3. cubefox · · focus · HN ↗
        > Model companies are doing this for themselves anyways

        No it's EU law.

        1. WD-42 · · focus · HN ↗
          Why can’t it be both?
        2. [deleted] · · focus · HN ↗

          [deleted]

      4. charcircuit · · focus · HN ↗
        Feeding synthetic data made from the model does not cause collapse. Anthropic would not care about distillation if it just caused people's models to suck.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.