{
  "id": 11356872,
  "title": "Gemini 4 Argon Wins 12 of 18 Benchmarks. Code Isn't One.",
  "url": "https://urgent.news/2026/10/02/gemini-4-argon-wins-12-of-18-benchmarks-code-isnt-one",
  "topic": "ai",
  "section": "AI",
  "published": "2026-10-02T04:41:14.000Z",
  "source": {
    "name": "Dev.to",
    "slug": "dev-to",
    "url": "https://dev.to/max_quimby/gemini-4-argon-wins-12-of-18-benchmarks-code-isnt-one-2h70"
  },
  "original_language": "en",
  "account": "Google unveiled Gemini 4 Argon, a new AI model that outperforms GPT-6 Astra, Astra, and Claude Opus 5.5 on 12 out of 18 benchmarks. Argon excels in reasoning, knowledge, and multimodal tasks, securing top spots in DeepSWE v1.1, LVBench video understanding, and LMArena Text Arena. However, it lags behind in code-related benchmarks, ranking 8th in Code Arena WebDev. Notably, Argon offers a 1-million-token output window, surpassing previous models' 64K token cap. Despite its impressive performance, Argon's cost per task is higher than GPT-6 Astra and Claude Opus 5.5, though it delivers superior token consumption efficiency.",
  "summary": "Gemini 4 Argon Wins 12 of 18 Benchmarks. Code Isn't One. Google just dropped Gemini 4 Argon , and the numbers are hard to argue with. The model leads 12 of 18 published benchmarks against GPT-6 Astra and Claude Opus 5.5 -- scoring 77.9% on DeepSWE v1.1, 91.7% on LVBench video understanding, and landing the number one slot on LMArena Text Arena with 1,525 points. It ships a 1-million-token output…",
  "key_points": [
    "Gemini 4 Argon outperforms GPT-6 Astra, Astra, and Claude Opus 5.5 on 12 of 18 benchmarks",
    "Argon excels in reasoning, knowledge, and multimodal tasks",
    "Argon lags behind in code-related benchmarks, ranking 8th in Code Arena WebDev"
  ],
  "editors_take": "Gemini 4 Argon's broad outperformance on benchmarks positions it as a strong contender in AI, but its high cost per task and weakness in code-related tasks limit its appeal.",
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}