Claude Opus 5 scored 30% on ARC-AGI-3. Wrapped in Nvidia’s AVO, it hit 100%.
Nvidia first introduced its Agentic Variation Operators (AVO) general-purpose coding agent system in late March 2026. The company has now The post Claude Opus 5 scored 30% on ARC-AGI-3. Wrapped in Nvidia’s AVO, it hit 100%. appeared first on The New Stack .
Nvidia unveiled its Agentic Variation Operators (AVO) general-purpose coding agent system in March 2026. A team of software engineers, machine learning specialists and AI research interns at Nvidia explained how AVO boosted Claude Opus 5 from a baseline score of 30.2% on ARC-AGI-3 benchmark to a perfect score of 100% when integrated with the AVO agent system.
AVO specializes in sustaining autonomous operation across extended, multistep tasks, handling tasks such as inspecting and editing code, running commands, consulting documentation and validating work through execution. The ARC-AGI-3 benchmark uses RHAE (Relative Human Action Efficiency) metric, which combines task completion with per-level action efficiency relative to human baselines.
When Claude Opus 5 ran at max reasoning effort, it scored 97.5% on ARC-AGI-1 and 90.4% on ARC-AGI-2. AVO achieves a perfect RHAE score of 100.00 across all 25 environments in ARC-AGI-3 public set, completing all 183 levels. AVO's unique approach involves replacing predefined variation steps in conventional evolutionary-search systems with an autonomous agent that decides how to generate the next candidate - what to inspect, change, test, and commit.
The Nvidia team noted that both GPU-kernel optimization and ARC-AGI-3 benchmark involve building hypotheses from incomplete evidence, taking actions through external interfaces, observing consequences, preserving state, revising models, and recovering from incorrect assumptions, all while progressing over long horizons. To achieve sustained autonomous progress, AVO incorporates persistent memory and supervision.
Persistent memory carries forward prior implementations, evaluation results, compiler and profiler outputs, and accumulated reasoning, allowing the agent to resume from current state. Supervision involves a programmatic software module within the AVO system that monitors the broader search trajectory and can intervene when progress stalls.
Nvidia AVO was demonstrated on difficult software engineering and GPU-kernel optimization tasks. While Claude Opus 5 was used in the public-set result, the system was also paired with GPT-5.6 Sol in challenging game scenarios, where Sol demonstrated faster wall-clock time but Opus used fewer environment actions to match levels.
Written by urgent.news from The New Stack's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.