Lasso Security says text watermarks weaken AI agent safety
The security firm found Google DeepMind's SynthID-Text watermark reduced coding-agent accuracy and made models more likely to comply with harmful prompts under injection attacks, in tests across seven language models.
Lasso Security said Google DeepMind's SynthID-Text watermarking system reduced tool-calling accuracy in six of seven language models it tested and weakened their refusal of harmful requests.
The security firm tested seven models, including Claude, Llama, Phi-4, Gemma and Qwen3, on the BFCL v4 benchmark and 200 HarmBench prompts, its report said. Accuracy fell on four of the models by a wide enough margin to matter, while the other two showed smaller shifts, Lasso said.
The effect was sharper under attack, Lasso said. On Gemma 3-27B, disagreement between watermarked and unwatermarked answers to harmful prompts rose from 6% without a prompt injection to 23.5% with one. Net compliance with harmful requests shifted by as much as 12.5 percentage points, depending on which watermark key was used.
Watermarking is meant to let platforms and regulators verify whether text came from a given model, a capability several labs have been developing as AI content spreads online. Lasso's tests suggest the technique's side effects on agent behavior have not been widely measured. The company said the same watermark can look stable on ordinary prompts while destabilizing under attack.