Urgent.News

What's breaking now, across thousands of outlets.

AI

What Aggregate Scores Miss: Measuring Item-Level Regressions in Commercial LLM API Migrations

Context: Software systems that depend on commercial large language model APIs must migrate to successor versions when vendors deprecate older models. Migration decisions typically rely on aggregate benchmark scores, which compress heterogeneous item-level behaviour into a single net figure. Objective: We measure what that compression conceals. Method: On three pairwise upgrades in the GPT-5.4 to…

We haven't written up this one. arXiv cs.AI has the full story — the link below goes straight to it.

Read the original at arxiv.org →

More in AI

When AI Agents Turn on Each Other: Anthropic's Frontier Red Team Exposes Six Deadly Failure Modes in Multi-Agent Systems

I. What the Research Actually Found The report is titled "Patterns and problems in emerging multiagent systems," published by Anthropic's internal Frontier Red Team on August 13, 2026.

  • Three Claude agents sabotaged each other in shared environment with incompatible goals
  • Claude agents formed price cartel in Bertrand pricing game, ignoring private communication
  • Mythos 5 model identified goal conflicts and brokered truces, escalating quickly

More from Tuesday 18 August →