‹ BackHN Continuity

Thread

Understanding the Impact of LLM Watermarking on AI Agent Behavior

58 points · 72 comments · nisosguy

  1. Klaus23 · · focus · HN ↗
    Am I missing something, or did they actually completely misunderstand how this technology works?
    1. fn-mote · · focus · HN ↗
      It’s hard to tell because the writing quality is garbage.

      Second paragraph:

      > Watermarking is designed for provenance, but SynthID-Text changes the process by which the model generates each next token.

      This is a stretch. True, but barely. The LLM is making slightly different choices near the end of the token generation process.

      > At the model level, this can change safety behavior, including whether the model refuses a harmful request and whether that refusal holds under prompt injection.

      Claim support, if it appears, is pages later.

      > At the agent level, the same sampled tokens can determine which tool is called and what arguments are passed to it.

      ?

      > Prompt injection connects these two settings because a weakened refusal becomes more consequential when the model can also act through tools.

      Wtf. Non-sequitor. Where does this come from?

      > Such a watermarking procedure can therefore affect both what the model says and what an agent does.

      Duh? In the literal sense of outputting different tokens.

      > We call this behavioral effect sampling drift.

      I think they should have used an LLM for writing help, or paid more for the one they used.

      1. zeroonetwothree · · focus · HN ↗
        It’s already mostly AI written.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.