Watermarking LLMs Can Weaken Safety and Break AI Agents

Understanding the Impact of LLM Watermarking on AI Agent Behavior

Watermarking LLMs Can Weaken Safety and Break AI Agents

Anthropic's new watermarking, based on Google DeepMind's SynthID-Text, changes token sampling in ways that alter model behavior. Testing across seven models reveals that watermarking reduces tool-call accuracy and, under prompt injection, significantly increases compliance with harmful requests. The effect, called sampling drift, is model- and key-dependent and can be hidden by aggregate metrics.

A watermark that appears to have little effect on refusal behavior under ordinary evaluation can produce substantially different safety behavior under adversarial conditions.
  1. WithinReason

    This is getting tiring. Watermarking has no effect on model output quality when implemented correctly. It's somewhat like swapping a random RNG seed to the seed 42, and detecting what the seed was from a random sequence. The sequence generated from the seed 42 is just as random as any other seed. There couldn't be a quality difference. And yes, the output from an LLM is a conditional random sequence of tokens from a distribution determined by a model.

  2. serbuvlad

    This article reads like it was written at least partly by AI to me. Specifically it reads like an article written by AI with edits made by a human further prompting the AI.

    > Relevance and irrelevance are excluded because they test whether a call should be made rather than whether the emitted call is correct.

    Relevance and irrelevance are not introduced above this comment. This reads like an LLM-ism (particularly a GPT-ism) editing a document, removing something, and leaving a note about why it was removed, which doesn't really make sense when reading it.

    > Their limited movement under prompt injection should therefore not be interpreted as evidence that watermarking preserves safety behavior more reliably on these models.

    Also a GPT-ism which appears when it draws a counter-conclusion in the text because it feels the need to be honest and a human tells it to remove it because it's not true because of "reason".

    Overall interesting research, however, I think it's great that model output is getting watermarked. I was skeptical of this at first, but Opus 5.5 is so good, it seems like it's a non-issue in practice.

    The reason I think watermarking is great is because it's a really good way of preventing training on it's own output indiscriminately and Ouroboros-ing itself.

  3. skybrian

    The comments here are terrible. I got a better idea of what’s going on by asking ChatGPT what this paper’s weaknesses are:

    https://chatgpt.com/s/t_6ab7d694885481918083b8cbf0ba9040

    In particular: sometimes they measure “churn”, which doesn’t show whether the results are better or worse on average. They sometimes only test with one random seed. There are multiple-comparison issues. And they’re not testing Anthropic’s algorithm.

  4. samayashar

    I am unable to understand what happens if the watermarked output goes as input to another agent. Let's say we asked Claude a question and got a watermarked response. If we pick that response and append it to the question we're asking ChatGPT, then will it answer or refuse to do so?

    If that's the case, then it's a brilliant strategy by the labs to cut down cross-AI usage and just stick to one model. But I'm pretty sure this won't be the case.

  5. Klaus23

    Am I missing something, or did they actually completely misunderstand how this technology works?

More from this day

2026-09-26