Understanding the Impact of LLM Watermarking on AI Agent Behavior
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Understanding the Impact of LLM Watermarking on AI Agent Behavior
Unofficial Hacker News client; not affiliated with Y Combinator.
serbuvlad · · focus · HN ↗
> Relevance and irrelevance are excluded because they test whether a call should be made rather than whether the emitted call is correct.
Relevance and irrelevance are not introduced above this comment. This reads like an LLM-ism (particularly a GPT-ism) editing a document, removing something, and leaving a note about why it was removed, which doesn't really make sense when reading it.
> Their limited movement under prompt injection should therefore not be interpreted as evidence that watermarking preserves safety behavior more reliably on these models.
Also a GPT-ism which appears when it draws a counter-conclusion in the text because it feels the need to be honest and a human tells it to remove it because it's not true because of "reason".
Overall interesting research, however, I think it's great that model output is getting watermarked. I was skeptical of this at first, but Opus 5.5 is so good, it seems like it's a non-issue in practice.
The reason I think watermarking is great is because it's a really good way of preventing training on it's own output indiscriminately and Ouroboros-ing itself.
zeroonetwothree · · focus · HN ↗
serbuvlad · · focus · HN ↗
I just took his text, pasted it to ChatGPT, "rephrase", paste it back in the online checker, 0% AI.
gruez · · focus · HN ↗
unsnap_biceps · · focus · HN ↗