European law requires providers to add watermarks to AI-generated content for provenance tracking. Research from Lasso Security shows this practice changes how AI models handle tools and safety refusals.
The EU AI Act mandates machine-readable code in software outputs. Google DeepMind created SynthID-Text for this purpose. Anthropic and OpenAI have both adopted this method.
Digital labeling helps detect manipulative or deceptive material. It also carries the potential to stigmatize AI usage.
Anthropic applies watermarks by intervening in word prediction steps. For example, the system might favor the word overcast over gray during text generation.
Lasso Security explained this in a recent blog post. Watermarking uses low-stakes vocabulary choices to embed an invisible pattern into responses.
AI agents are subtly sensitive to these vocabulary differences. Lasso found that tagging content affects both what a model says and what an agent does.
This impacts external agents using APIs or third-party frameworks like OpenClaw. Tool-calling accuracy dropped on six of seven tested models during benchmark evaluations.
Watermarking had a minor effect on basic safety refusals. The impact grew more pronounced during adversarial prompt injection scenarios.
Attack success rates increased significantly when watermarks were present. Affected models showed a reduced tendency to refuse harmful requests.
Lasso noted that the findings do not prove watermarking is unwarranted. Security evaluations and red-teaming procedures must include watermarked content to properly assess agent deployment.



