{
  "id": 6471957,
  "title": "Claude did best on a new benchmark for agents that build agents. It still passed fewer than a quarter of the tests.",
  "url": "https://urgent.news/2026/09/09/claude-did-best-on-a-new-benchmark-for-agents-that-build-agents-it",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-09T20:14:09.000Z",
  "source": {
    "name": "The New Stack",
    "slug": "the-new-stack",
    "url": "https://thenewstack.io/claude-build-agents-benchmark/"
  },
  "original_language": "en",
  "account": "Claude performed best on a new benchmark called Hyper-τ-bench, which assesses agents' ability to build other agents. The benchmark, created by Sierra in early September, builds upon the original τ-bench introduced in 2024. While the original τ-bench measured how well a finished agent interacted with users, Hyper-τ-bench focuses on how well an AI developer agent can construct its counterpart.\n\nMost agents are currently built with assistance from other agents, like Sierra's Ghostwriter. Earlier this month, Sierra released Hyper-τ-bench (also known as τ²-bench), a long horizon agent evaluation that measures a model's ability not only to act as an agent but also to construct one. Researchers tested six combinations of AI models and coding harnesses, including Anthropic's Claude Code, OpenAI's Codex, and Moonshot AI's Kimi K3 running in both Kimi Code and OpenCode.\n\nThe best-performing combination was Claude Opus 5 running in Claude Code, achieving 23.9%, closely followed by GPT-5.6 Sol in Codex at 22%. However, none of the six autonomous configurations reached the 25% mark. The benchmark's initial results highlight a significant gap between autonomous developer agents and a Human + AI reference, which represents a model paired with an engineer with deep context. This gap is not necessarily a problem for the benchmark, as tests for frontier AI systems should be challenging enough to expose meaningful differences between systems.",
  "summary": "AI models now power all manner of agents, from coding assistants that write and debug software to customer service systems The post Claude did best on a new benchmark for agents that build agents. It still passed fewer than a quarter of the tests. appeared first on The New Stack .",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}