{
  "id": 7209810,
  "title": "Two \"Codex CLI\" models on the same benchmark: the harness hides the model",
  "url": "https://urgent.news/2026/09/14/two-codex-cli-models-on-the-same-benchmark-the-harness-hides-the-model",
  "topic": "tech",
  "section": "Tech",
  "published": "2026-09-14T00:15:06.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/cole_halton_42f71d71b809b/two-codex-cli-models-on-the-same-benchmark-the-harness-hides-the-model-1jn"
  },
  "original_language": "en",
  "account": "Two Codex CLI models on the same benchmark demonstrate the critical importance of considering both the model and its underlying harness. Specific Labs released Real-SWE, an enterprise-code SWE benchmark, and the leaderboard reveals a striking 2x gap between GPT-6 Astra on Codex CLI (33.8% resolution) and GPT-5.6 Sol on Codex CLI (16.2% resolution). This emphasizes that the harness, or routing layer, plays a significant role in performance, not just the underlying model. Real-SWE's approach of framing results as model-and-harness combinations rather than isolated models is a more honest way to present data. However, many vendors fail to disclose this detail, as a lower number under their own tooling looks unfavorable. This highlights the need for readers to carefully examine both the model and harness when evaluating coding agents, as the route (harness) is often more influential than the brain (model) itself.",
  "summary": "Specific Labs dropped Real-SWE, an enterprise-code SWE benchmark, and the leaderboard is a great study in why you should never read \"Claude Code\" or \"Codex CLI\" as a model name. Same harness, two different brains: GPT-6 Astra on Codex CLI: 33.8% resolution GPT-5.6 Sol on Codex CLI: 16.2% resolution Same vendor's CLI, same harness, same benchmark. More than a 2x gap. If you'd just read \"Codex CLI…",
  "key_points": [],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 2,
    "also_reported_by": [
      {
        "outlet": "Dev.to",
        "title": "How to change reasoning effort in Codex CLI: model_reasoning_effort values and one-off overrides",
        "url": "https://urgent.news/2026/09/13/how-to-change-reasoning-effort-in-codex-cli-model-reasoning-effort",
        "published": "2026-09-13T15:25:00.000Z"
      }
    ]
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}