‹ BackHN Continuity

Thread

Understanding the Impact of LLM Watermarking on AI Agent Behavior

58 points · 72 comments · nisosguy

  1. serbuvlad · · focus · HN ↗
    This article reads like it was written at least partly by AI to me. Specifically it reads like an article written by AI with edits made by a human further prompting the AI.

    > Relevance and irrelevance are excluded because they test whether a call should be made rather than whether the emitted call is correct.

    Relevance and irrelevance are not introduced above this comment. This reads like an LLM-ism (particularly a GPT-ism) editing a document, removing something, and leaving a note about why it was removed, which doesn't really make sense when reading it.

    > Their limited movement under prompt injection should therefore not be interpreted as evidence that watermarking preserves safety behavior more reliably on these models.

    Also a GPT-ism which appears when it draws a counter-conclusion in the text because it feels the need to be honest and a human tells it to remove it because it's not true because of "reason".

    Overall interesting research, however, I think it's great that model output is getting watermarked. I was skeptical of this at first, but Opus 5.5 is so good, it seems like it's a non-issue in practice.

    The reason I think watermarking is great is because it's a really good way of preventing training on it's own output indiscriminately and Ouroboros-ing itself.

    1. zeroonetwothree · · focus · HN ↗
      Pangram flags it as mostly AI.
      1. serbuvlad · · focus · HN ↗
        I don't trust those. My friend in med school had an issue last year because he was writing his stuff himself and he was getting flagged as AI in this online checker (which his professors used to check submissions). He asked me about it.

        I just took his text, pasted it to ChatGPT, "rephrase", paste it back in the online checker, 0% AI.

        1. gruez · · focus · HN ↗
          The burden of proof required for accusing someone of academic misconduct should surely be different than accusing someone of AI slop?
          1. unsnap_biceps · · focus · HN ↗
            There's been a ton of news stories of people claiming otherwise.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.