Urgent.News

What's breaking now, across thousands of outlets.

AI

Claude did best on a new benchmark for agents that build agents. It still passed fewer than a quarter of the tests.

AI models now power all manner of agents, from coding assistants that write and debug software to customer service systems The post Claude did best on a new benchmark for agents that build agents. It still passed fewer than a quarter of the tests. appeared first on The New Stack .

Claude did best on a new benchmark for agents that build agents. It still passed fewer than a quarter of the tests.

Claude performed best on a new benchmark called Hyper-τ-bench, which assesses agents' ability to build other agents. The benchmark, created by Sierra in early September, builds upon the original τ-bench introduced in 2024. While the original τ-bench measured how well a finished agent interacted with users, Hyper-τ-bench focuses on how well an AI developer agent can construct its counterpart.

Most agents are currently built with assistance from other agents, like Sierra's Ghostwriter. Earlier this month, Sierra released Hyper-τ-bench (also known as τ²-bench), a long horizon agent evaluation that measures a model's ability not only to act as an agent but also to construct one. Researchers tested six combinations of AI models and coding harnesses, including Anthropic's Claude Code, OpenAI's Codex, and Moonshot AI's Kimi K3 running in both Kimi Code and OpenCode.

The best-performing combination was Claude Opus 5 running in Claude Code, achieving 23.9%, closely followed by GPT-5.6 Sol in Codex at 22%. However, none of the six autonomous configurations reached the 25% mark. The benchmark's initial results highlight a significant gap between autonomous developer agents and a Human + AI reference, which represents a model paired with an engineer with deep context.

This gap is not necessarily a problem for the benchmark, as tests for frontier AI systems should be challenging enough to expose meaningful differences between systems.

Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at thenewstack.io →

More in AI

Google Revamps AI Subscription Plans With New Ultra Tiers and Gmail Productivity Tools

Google has expanded its AI subscription lineup with new Ultra options, broader access to its latest models and productivity features that connect Gemini app more closely with Gmail, Calendar and other…

  • Google revamps AI subscription plans with new Ultra tiers.
  • Adds Gmail productivity tools like AI Inbox and Daily Brief.
  • Shifts subscription model from daily prompt caps to usage-based.

More from Wednesday 9 September →