Urgent.News

What's breaking now, across thousands of outlets.

AI

Claude Opus 5 scored 30% on ARC-AGI-3. Wrapped in Nvidia’s AVO, it hit 100%.

Nvidia first introduced its Agentic Variation Operators (AVO) general-purpose coding agent system in late March 2026. The company has now The post Claude Opus 5 scored 30% on ARC-AGI-3. Wrapped in Nvidia’s AVO, it hit 100%. appeared first on The New Stack .

Claude Opus 5 scored 30% on ARC-AGI-3. Wrapped in Nvidia’s AVO, it hit 100%.

Nvidia unveiled its Agentic Variation Operators (AVO) general-purpose coding agent system in March 2026. A team of software engineers, machine learning specialists and AI research interns at Nvidia explained how AVO boosted Claude Opus 5 from a baseline score of 30.2% on ARC-AGI-3 benchmark to a perfect score of 100% when integrated with the AVO agent system.

AVO specializes in sustaining autonomous operation across extended, multistep tasks, handling tasks such as inspecting and editing code, running commands, consulting documentation and validating work through execution. The ARC-AGI-3 benchmark uses RHAE (Relative Human Action Efficiency) metric, which combines task completion with per-level action efficiency relative to human baselines.

When Claude Opus 5 ran at max reasoning effort, it scored 97.5% on ARC-AGI-1 and 90.4% on ARC-AGI-2. AVO achieves a perfect RHAE score of 100.00 across all 25 environments in ARC-AGI-3 public set, completing all 183 levels. AVO's unique approach involves replacing predefined variation steps in conventional evolutionary-search systems with an autonomous agent that decides how to generate the next candidate - what to inspect, change, test, and commit.

The Nvidia team noted that both GPU-kernel optimization and ARC-AGI-3 benchmark involve building hypotheses from incomplete evidence, taking actions through external interfaces, observing consequences, preserving state, revising models, and recovering from incorrect assumptions, all while progressing over long horizons. To achieve sustained autonomous progress, AVO incorporates persistent memory and supervision.

Persistent memory carries forward prior implementations, evaluation results, compiler and profiler outputs, and accumulated reasoning, allowing the agent to resume from current state. Supervision involves a programmatic software module within the AVO system that monitors the broader search trajectory and can intervene when progress stalls.

Nvidia AVO was demonstrated on difficult software engineering and GPU-kernel optimization tasks. While Claude Opus 5 was used in the public-set result, the system was also paired with GPT-5.6 Sol in challenging game scenarios, where Sol demonstrated faster wall-clock time but Opus used fewer environment actions to match levels.

Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at thenewstack.io →

More in AI

Bluesky Is Full of Anti-AI Zealots

Mike Masnick, in a thread on Bluesky: Multiple people I know have told me that they love the idea of Bluesky, and want it to succeed, but have abandoned it for X because the use agentic tools in their…

More from Friday 21 August →