Understanding the Impact of LLM Watermarking on AI Agent Behavior
Lasso Security finds SynthID-Text watermarking can alter LLM refusals and agent tool calls through sampling drift.
Lasso Security reports that Anthropic's planned Claude watermark, based on Google DeepMind's SynthID-Text, can change behavior even in a non-distortionary configuration. Paired tests on BFCL v4 tool calling and HarmBench refusals, including a fixed prompt-injection setup, found model- and key-dependent disagreement between watermarked and unwatermarked runs. The authors say this sampling drift can change which tool an agent calls and the arguments it passes, and whether refusals of harmful requests hold. They also note EU AI Act Article 50(2) requires machine-readable marking of synthetic text.