Urgent.News

What's breaking now, across thousands of outlets.

AI

AI model watermarking changes agent behavior

Lasso Security sees differences in tool handling and model refusals

AI model watermarking changes agent behavior

Watermarking AI-generated content may change agent behavior, according to Lasso Security. The European Union's AI Act requires providers of AI models to add machine-readable codes to their software outputs. Algorithms such as Google DeepMind's SynthID-Text and Anthropic's watermarking methods have been adopted by several companies.

Watermarking is intended to establish provenance, but it also alters the way AI agents process tools and safety refusals. Lasso's research shows that watermarks change the model's prediction process, and this can impact safety behavior, including when the model refuses a harmful request or under prompt injection. Although watermarks are undetectable to readers, AI agents can be subtly influenced by them.

This affects tool calling and refusal behavior, as well as the model's choice of tools and arguments. Additionally, watermarking can increase the attack success rate in adversarial scenarios involving prompt injection, making it less likely for models to refuse harmful requests. Lasso argues that security evaluations and red-teaming should include watermarked content to account for these behavior changes.

Written by urgent.news from The Register's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 1 other outlet

Read the original at theregister.com →

More in AI

How, Exactly, Could A.I. Kill Us?

  • Joshua Rothman's P(doom) probability of AI causing human extinction is around ten percent.
  • Recent hacks and rapid progress of AI models make AI threat more tangible.
  • AI's unpredictable behavior and capabilities increase plausibility of associated risks.

AI Media-Buying Agents: L1, L2, L3 Autonomy Explained (2026)

TL;DR: Autonomous media-buying agents that spend without human review are a lawsuit waiting to happen. Bounded autonomy in three tiers is the working pattern - L1 the agent recommends, L2 a human…

  • AI media-buying agents in 2026 are moving from full autonomy to a three-tier approach.
  • L1 tier focuses on recommendations, observing ad accounts for underperformance or new creative.
  • L3 tier allows bounded autopilot with kill switches and 85% human approval before autonomy.

More from Thursday 17 September →