‹ BackHN Continuity

Thread

Understanding the Impact of LLM Watermarking on AI Agent Behavior

58 points · 72 comments · nisosguy

  1. willmadden · · focus · HN ↗
    Watermarking sounds like a good idea, but it's not. Token drift from watermarking will degrade the quality of outputs and could allow clever people to circumvent guardrails.

    You should assume all text is AI generated. If you want to "test" someone at school or during an interview, have them write with a pencil and paper.

    1. [deleted] · · focus · HN ↗

      [deleted]

    2. skybrian · · focus · HN ↗
      Changing a random seed could either improve or degrade the output. In theory, better and worse outputs should be equally probable, depending on your luck.
      1. riskable · · focus · HN ↗
        They tested this in the research: They found that the statistical likelihood of bad tool calling was higher with watermarking compared to random seeds. At least, that was my takeaway (I read the whole thing).

        This is because SynthID and similar watermarking methods for LLMs don't just change the random seed. They take additional steps (that I have yet to read about) in order to detect when someone changes a few words of the output, trying to remove the watermark.

        The other takeaway is that just by knowing a watermark is being applied gives an attacker an advantage in working around safety features because then they know the output isn't based on true randomness and can take advantage of that in a similar fashion to how breaking cryptography becomes easier when the RNG isn't truly random.

    3. possibilistic · · focus · HN ↗
      It's a horrible idea. These companies can embed unique identifiers in content to forever track you and your content's diffusion across the web:

      <a href="https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49794615">https:&#x2F;&#x2F;news.ycombinator.com&#x2F;item?id=49794615

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.