Urgent.News

What's breaking now, across thousands of outlets.

Tech

Two "Codex CLI" models on the same benchmark: the harness hides the model

Specific Labs dropped Real-SWE, an enterprise-code SWE benchmark, and the leaderboard is a great study in why you should never read "Claude Code" or "Codex CLI" as a model name. Same harness, two different brains: GPT-6 Astra on Codex CLI: 33.8% resolution GPT-5.6 Sol on Codex CLI: 16.2% resolution Same vendor's CLI, same harness, same benchmark. More than a 2x gap. If you'd just read "Codex CLI…

Two Codex CLI models on the same benchmark demonstrate the critical importance of considering both the model and its underlying harness. Specific Labs released Real-SWE, an enterprise-code SWE benchmark, and the leaderboard reveals a striking 2x gap between GPT-6 Astra on Codex CLI (33.8% resolution) and GPT-5.6 Sol on Codex CLI (16.2% resolution).

This emphasizes that the harness, or routing layer, plays a significant role in performance, not just the underlying model. Real-SWE's approach of framing results as model-and-harness combinations rather than isolated models is a more honest way to present data. However, many vendors fail to disclose this detail, as a lower number under their own tooling looks unfavorable.

This highlights the need for readers to carefully examine both the model and harness when evaluating coding agents, as the route (harness) is often more influential than the brain (model) itself.

Written by urgent.news from Dev.to's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

This story

This is one outlet's version. Read the fullest account.

Read the original at dev.to →

More in Tech

WE:=DAEMONCORE_ACADAMY

DaemonCore Academy: Cybersecurity Education Should Not Have a Cover Charge We built DaemonCore Academy around a fairly radical idea. Education should be free.

More from Monday 14 September →