Urgent.News

What's breaking now, across thousands of outlets.

AI

GPT-6 Astra scores 62.7% on ARC-AGI-3 with the standard harness and 99.9% with a new provider adapter harness; Claude Opus 5 scored 30.2%, and GPT-5.6 Sol 7.8% (Greg Kamradt/ARC Prize)

GPT-6 Astra scores 62.7% for $26K on ARC-AGI-3 Semi-Private with our Standard harness, and 99.9% for $19K with a Provider Adapter harness.

OpenAI's Astra achieved a remarkable 62.7% score on ARC-AGI-3 using the Standard harness, while utilizing a new Provider Adapter harness, Astra reached an impressive 99.9% score. In comparison, Claude Opus 5 obtained a lower 30.2% score and GPT-5.6 Sol only managed 7.8%. These scores demonstrate the significant advancements in artificial intelligence capabilities and the growing gap between AI and human performance.

Written by urgent.news from Techmeme's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Also reported by 2 other outlets

Read the original at arcprize.org →

More in AI

Getting Agents to Stop Assuming: What a First AWS Agent Workflow Reveals About Constraint Design

The most common failure mode in agent workflows is not a timeout or a bad API call. It is the agent confidently doing the wrong thing because it filled in missing information with a plausible guess.

  • Agents often create incorrect results by filling missing information with plausible guesses.
  • Validation gates in AWS Bedrock's AgentCore prevent agents from proceeding with incomplete data.

More from Thursday 3 September →