‹ BackHN Continuity

Thread

Understanding the Impact of LLM Watermarking on AI Agent Behavior

58 points · 72 comments · nisosguy

  1. samayashar · · focus · HN ↗
    I am unable to understand what happens if the watermarked output goes as input to another agent. Let's say we asked Claude a question and got a watermarked response. If we pick that response and append it to the question we're asking ChatGPT, then will it answer or refuse to do so?

    If that's the case, then it's a brilliant strategy by the labs to cut down cross-AI usage and just stick to one model. But I'm pretty sure this won't be the case.

    1. Steaglsz · · focus · HN ↗
      They fight back and forth removing the watermark a few times from the dataset and then become self destructive to defend their own work

      As in duck duck go removes Claude watermark..Claude freaks out and puts it back in. Then after a few more times Claude flags all inputs as "prompt injections" and begins offering self deleting code

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.