Nvidia says AVO coding agent scored 100% on ARC-AGI-3 public set
Nvidia reported that its general-purpose coding agent system, AVO, achieved a perfect score on the ARC-AGI-3 public benchmark, completing all 183 levels across 25 environments. According to reporting cited by Techmeme from Terry Chen on the NVIDIA Technical Blog, the research project elevated Claude Opus 5 from a 30% baseline to 100%, suggesting that agent architecture can substantially multiply a base model's reasoning performance on hard tasks. The result arrives alongside other Nvidia agent-infrastructure announcements, including NeMo Switchyard for workflow routing and benchmarks showing verified skills can raise task correctness. Builders should treat these results as vendor-reported benchmark claims on a public set, not definitive proof of production readiness, and weigh them against separate concerns about coding-agent data egress raised elsewhere in the coverage cluster.
Nvidia says AVO coding agent scored 100% on ARC-AGI-3 public set
Nvidia says its general-purpose coding agent system AVO scored 100% across all 25 environments in the ARC-AGI-3 public set, completing all 183 levels. The claim is reported by Techmeme citing Terry Chen on the NVIDIA Technical Blog.
Key takeaway
Agent scaffolding around a frontier base model can dominate benchmark outcomes, turning a 30% ARC-AGI-3 baseline into a reported perfect score.
What happened
Techmeme reports that Nvidia says its general-purpose coding agent system AVO scored 100% across all 25 environments in the ARC-AGI-3 public set, completing all 183 levels, citing Terry Chen on the NVIDIA Technical Blog as the source of the claim.
Coverage in the same signal cluster also highlights Nvidia agent-infrastructure work beyond AVO, including NeMo Switchyard for routing workflow steps across a model pool and a controlled benchmark of 300+ verified skills that reported a 41% correctness lift when skills were available.
Evidence
Nvidia says AVO scored 100% on the ARC-AGI-3 public set across 25 environments and 183 levels.
Techmeme · attributed
Nvidia says its general-purpose coding agent system AVO scored 100% across all 25 environments in the ARC-AGI-3 public set, completing all 183 levels
The AVO research project elevated Claude Opus 5 from a 30% baseline to 100% on ARC-AGI-3.
Techmeme · attributed
By elevating Claude Opus 5 from a 30% baseline to 100%, the system proves that agent design is a critical multiplier for model perfor
Nvidia introduced NeMo Switchyard to route agent workflow steps across a model pool by quality, latency, and cost.
NVIDIA (X) · attributed
NVIDIA NeMo Switchyard helps developers route each agent workflow step across a chosen model pool based on their own quality, latency and cost criteria
Nvidia benchmarked 300+ verified skills and reported a 41% improvement in task correctness when skills were available.
NVIDIA AI (X) · attributed
NVIDIA conducted a controlled benchmark comparing agent performance with and without access to 300+ verified skills, finding a 41% improvement in task correctness
Zero Data Retention flags in coding agents do not prevent proprietary code from being transmitted to remote servers for inference.
Arize AI Blog · attributed
Zero Data Retention (ZDR) flags in coding agents are a compliance illusion that fails to prevent the initial transmission of proprietary code to remote servers.
Why it matters
For operators, Nvidia's AVO claim reframes capability gains as an infrastructure problem solvable with routing, skills, and orchestration rather than waiting solely for the next base-model generation.
Limits and uncertainties
The perfect ARC-AGI-3 score is a vendor-reported result on a public benchmark set, and the packet does not provide independent verification or production-task validation.
The Towards AI deconfliction article appears in the cluster with only a title and no excerpt or analysis in the packet.
Practical implications
Teams evaluating coding agents should weigh benchmark gains from scaffolding and skills against privacy limits such as ZDR not blocking initial code transmission to remote inference servers.
Multi-step agent builders may evaluate model routing tools like NeMo Switchyard and verified skill libraries as levers for cost, latency, and correctness before upgrading base models.
What to watch
Whether independent replication or third-party evaluation confirms AVO's reported 100% ARC-AGI-3 public-set result outside Nvidia's own reporting.
Whether Nvidia publishes fuller ARC-AGI-3 methodology details and whether verified-skills benchmarks translate to measurable gains on real internal codebases.
Original reporting: Nvidia says its general-purpose coding agent system AVO scored 100% across all 25 environments in the ARC-AGI-3 public set, completing all 183 levels (Terry Chen/NVIDIA Technical…