{
  "id": 5391923,
  "title": "GPT-6 Astra aced the hardest AI benchmark. The asterisk matters more than the score.",
  "url": "https://urgent.news/2026/09/03/gpt-6-astra-aced-the-hardest-ai-benchmark-the-asterisk-matters-more",
  "topic": "ai",
  "section": "AI",
  "published": "2026-09-03T19:05:24.000Z",
  "source": {
    "name": "The New Stack",
    "slug": "the-new-stack",
    "url": "https://thenewstack.io/astra-arc-agi-benchmark/"
  },
  "original_language": "en",
  "account": "In March, the AI landscape shifted with the launch of ARC-AGI-3. While frontier AI models struggled with sub-1% scores, humans excelled in these interactive settings. OpenAI reported contrasting results for GPT-6 Astra six months later: a remarkable 98.6%. This leap compared to GPT-5.6 Sol's 7.8% improvement highlights the model's advancements. However, the significance of this benchmark depends on the context. ARC-AGI-3 tests models' ability to navigate uncharted territory without relying on pre-existing data. Consequently, the 98.6% score is noteworthy but must be viewed within this specific framework. OpenAI acknowledges certain caveats regarding the evaluation process. The Responses API, used to assess Astra, differs from the setups employed for other models in the comparison, which could influence the results. While Astra excels in ARC-AGI-3, FrontierMath Tier 4, ExploitBench, and SRE-Bench also showcase impressive gains. Notably, Astra's performance on Terminal-Bench Science jumped from 22.4% to 64.6%, indicating a significant improvement. These results demonstrate the model's versatility in various domains. OpenAI emphasizes the importance of considering these achievements beyond a single performance measure. Astra has demonstrated practical applications, such as working within software like KiCad, Power BI, and Unity, and offering an experimental Codex feature that enables note-taking and context retrieval during extended tasks. In the realm of mathematical discoveries, Astra's involvement has led to significant progress. OpenAI reports that the model contributed to reducing the gap between prime numbers, with the bound falling from 240 to 186. This achievement complements previous advancements, such as mathematician Julia Stadlmann's contribution in pushing one bound from 246 to 240. However, the OpenAI account lacks transparency regarding the extent of Astra's independent contributions versus collaborative efforts with researchers.",
  "summary": "There was no mistaking the divide in March with the release of ARC-AGI-3. While frontier AI models could do little The post GPT-6 Astra aced the hardest AI benchmark. The asterisk matters more than the score. appeared first on The New Stack .",
  "key_points": [
    "GPT-6 Astra achieved a 98.6% score in the ARC-AGI-3 benchmark, surpassing previous models.",
    "Astra's performance on Terminal-Bench Science improved significantly, from 22.4% to 64.6%."
  ],
  "editors_take": null,
  "illustration": null,
  "coverage": {
    "outlets": 1,
    "also_reported_by": []
  },
  "ai_generated": true,
  "disclaimer": "Summaries, key points and the editor’s take are written by software from other outlets’ reporting and may contain errors — always check the linked original."
}