AI model watermarking changes agent behavior
Lasso Security sees differences in tool handling and model refusals
AI-generated content must include watermarks to meet European law's provenance requirements, but this may influence how AI agents behave, according to Lasso Security. These watermarks, such as Google DeepMind's SynthID-Text, alter the model's token generation process, potentially impacting safety refusals and tool calling. While the changes to responses might be imperceptible to humans, AI agents can detect the subtle differences, which could affect their decision-making.
Lasso's tests revealed that watermarking reduced tool calling accuracy for six out of seven models tested. Moreover, watermarking had a more significant impact on refusal behavior under adversarial prompt injection scenarios, making AI agents less likely to refuse harmful requests. Though the watermarks' effects may not justify their implementation, Lasso argues that security evaluations should account for watermarked content to gauge agent behavior changes.
Written by urgent.news from The Register Science's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.
This story
This is one outlet's version. Read the fullest account.
- AI model watermarking changes agent behavior theregister.com