Urgent.News

What's breaking now, across thousands of outlets.

AI

AI model watermarking changes agent behavior

Lasso Security sees differences in tool handling and model refusals

AI model watermarking changes agent behavior

AI-generated content must include watermarks to meet European law's provenance requirements, but this may influence how AI agents behave, according to Lasso Security. These watermarks, such as Google DeepMind's SynthID-Text, alter the model's token generation process, potentially impacting safety refusals and tool calling. While the changes to responses might be imperceptible to humans, AI agents can detect the subtle differences, which could affect their decision-making.

Lasso's tests revealed that watermarking reduced tool calling accuracy for six out of seven models tested. Moreover, watermarking had a more significant impact on refusal behavior under adversarial prompt injection scenarios, making AI agents less likely to refuse harmful requests. Though the watermarks' effects may not justify their implementation, Lasso argues that security evaluations should account for watermarked content to gauge agent behavior changes.

Written by urgent.news from The Register Science's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at theregister.com →

More in AI

How, Exactly, Could A.I. Kill Us?

  • Joshua Rothman's P(doom) probability of AI causing human extinction is around ten percent.
  • Recent hacks and rapid progress of AI models make AI threat more tangible.
  • AI's unpredictable behavior and capabilities increase plausibility of associated risks.

AI Media-Buying Agents: L1, L2, L3 Autonomy Explained (2026)

TL;DR: Autonomous media-buying agents that spend without human review are a lawsuit waiting to happen. Bounded autonomy in three tiers is the working pattern - L1 the agent recommends, L2 a human…

  • AI media-buying agents in 2026 are moving from full autonomy to a three-tier approach.
  • L1 tier focuses on recommendations, observing ad accounts for underperformance or new creative.
  • L3 tier allows bounded autopilot with kill switches and 85% human approval before autonomy.

More from Thursday 17 September →