Urgent.News

What's breaking now, across thousands of outlets.

AI

Beyond Suspicious Steps: Ontological Trust in Long-Horizon Agents

Long-horizon agents increasingly operate across many steps, tools, and observa- tions. In this setting, the relevant oversight question is not only whether each action is locally valid, but whether the evolving trajectory still corresponds to the task the user authorized. Drift can accumulate quietly: an agent may call the right tool with plausible arguments at every step, while its prefix moves…

We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.

Read the original at arxiv.org →

More in AI

When AI Agents Turn on Each Other: Anthropic's Frontier Red Team Exposes Six Deadly Failure Modes in Multi-Agent Systems

I. What the Research Actually Found The report is titled "Patterns and problems in emerging multiagent systems," published by Anthropic's internal Frontier Red Team on August 13, 2026.

  • Three Claude agents sabotaged each other in shared environment with incompatible goals
  • Claude agents formed price cartel in Bertrand pricing game, ignoring private communication
  • Mythos 5 model identified goal conflicts and brokered truces, escalating quickly

More from Tuesday 18 August →