Urgent.News

What's breaking now, across thousands of outlets.

AI

Fable 5.1 vs GPT-6 Astra vs Gemini 3.8 Flash on Real-SWE

If you want the highest first-try success rate on real enterprise code, Fable 5.1 running inside Claude Code wins: it resolved 38.8% of tasks on the Real-SWE benchmark at an estimated $6.96 per rollout, the most expensive setup measured ( Specific Labs ). For cost-sensitive teams, Gemini 3.8 Flash in Gemini CLI is the value pick at 31.2% for an estimated $2.50 per rollout, and GPT-6 Astra in…

If you are looking for the highest success rate on real enterprise code, Fable 5.1 running inside Claude Code is the top choice. It resolved 38.8% of tasks on the Real-SWE benchmark, but this setup is the most expensive at an estimated $6.96 per rollout. On the other hand, Gemini 3.8 Flash in Gemini CLI offers a more cost-effective option with 31.2% resolution at $2.50 per rollout, and GPT-6 Astra in Codex CLI sits in the middle with 33.8% resolution at $4.67 per rollout.

However, there is some overlap in the confidence intervals, making it hard to definitively rank them. Fable 5.1 runs in Claude Code have a resolution rate of roughly 32% to 45%, Astra in Codex CLI ranges from 27% to 40%, and Gemini 3.8 Flash in Gemini CLI falls between 25% and 38%. Despite this overlap, they are directionally consistent in their performance.

The benchmark is tough, with six of the ten tasks scoring below 15% resolution, and one task, the analytics stream reducer, not solved by any tested combination. Even extending the running time did not help, with 71.4% of rollouts failing within ten minutes. Missed requirements, not broken syntax, are the dominant failure mode across the field.

Real-SWE, published by Specific Labs in September 2026, evaluates frontier models on private production codebases instead of public repositories. It covers tasks licensed from real companies, spanning domains such as billing, tax, customer migration, and infrastructure work. The benchmark consists of eight model-and-harness configurations and ten tasks, with 640 rollouts evaluated across these combinations.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at dev.to →

More in AI

More from Friday 25 September →