Urgent.News

What's breaking now, across thousands of outlets.

AI

AI’s best coding agent fails 60% of the time — and the data backs it up

Claude Fable 5.1 just won a new coding benchmark despite failing more than six out of 10 times. Its 38.8% The post AI’s best coding agent fails 60% of the time — and the data backs it up appeared first on The New Stack .

AI’s best coding agent fails 60% of the time — and the data backs it up

Claude Fable 5.1 achieved a coding benchmark score of 38.8%, but it failed more than half of the time. Based on Real-SWE, a benchmark by Y Combinator-backed Specific Labs, the agent was dropped into private codebases from real companies and required solving problems similar to those engineers face daily. The code and its solutions remain private, making it less likely that the model encountered any of the code during training.

Fable 5.1, running via Claude Code, led the pack with a 38.8% score, followed by GPT-6 Astra on Codex CLI at 33.8% and Gemini 3.8 Flash on Gemini CLI at 31.2%. Six tasks proved too challenging for all models, with Real-SWE testing each model's performance using its own coding tool. Notably, six out of ten tasks recorded success rates below 15%.

Not a single model successfully completed the analytics stream reducer task across 64 attempts. Fable 5.1's most common errors were missing requirements (36.7%) and integration errors (34.7%), while Astra's errors were equally split between integration errors and unverified assumptions. The data suggests that coding problem-solving differs significantly from navigating unfamiliar production codebases.

Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at thenewstack.io →

More in AI

Power Play

AI needs more power, faster than America has ever built it. An investigation into the bottlenecks, the fixes, and the economics of AI development.

More from Monday 14 September →