Skip to main content
LLMgram · AI News · 2026-08-22

Nvidia says AVO coding agent scored 100% on ARC-AGI-3 public set

Nvidia says AVO coding agent scored 100% on ARC-AGI-3 public set

Nvidia reported that its general-purpose coding agent system, AVO, achieved a perfect score on the ARC-AGI-3 public benchmark, completing all 183 levels across 25 environments. According to reporting cited by Techmeme from Terry Chen on the NVIDIA Technical Blog, the research project elevated Claude Opus 5 from a 30% baseline to 100%, suggesting that agent architecture can substantially multiply a base model's reasoning performance on hard tasks. The result arrives alongside other Nvidia agent-infrastructure announcements, including NeMo Switchyard for workflow routing and benchmarks showing verified skills can raise task correctness. Builders should treat these results as vendor-reported benchmark claims on a public set, not definitive proof of production readiness, and weigh them against separate concerns about coding-agent data egress raised elsewhere in the coverage cluster.

Sources

Nvidia says AVO coding agent scored 100% on ARC-AGI-3 public set

Nvidia says AVO coding agent scored 100% on ARC-AGI-3 public set

Nvidia says its general-purpose coding agent system AVO scored 100% across all 25 environments in the ARC-AGI-3 public set, completing all 183 levels. The claim is reported by Techmeme citing Terry Chen on the NVIDIA Technical Blog.

Key takeaway

Agent scaffolding around a frontier base model can dominate benchmark outcomes, turning a 30% ARC-AGI-3 baseline into a reported perfect score.

What happened

Techmeme reports that Nvidia says its general-purpose coding agent system AVO scored 100% across all 25 environments in the ARC-AGI-3 public set, completing all 183 levels, citing Terry Chen on the NVIDIA Technical Blog as the source of the claim.

Coverage in the same signal cluster also highlights Nvidia agent-infrastructure work beyond AVO, including NeMo Switchyard for routing workflow steps across a model pool and a controlled benchmark of 300+ verified skills that reported a 41% correctness lift when skills were available.

Evidence

  • Nvidia says AVO scored 100% on the ARC-AGI-3 public set across 25 environments and 183 levels.

    Techmeme · attributed

    Nvidia says its general-purpose coding agent system AVO scored 100% across all 25 environments in the ARC-AGI-3 public set, completing all 183 levels

  • The AVO research project elevated Claude Opus 5 from a 30% baseline to 100% on ARC-AGI-3.

    Techmeme · attributed

    By elevating Claude Opus 5 from a 30% baseline to 100%, the system proves that agent design is a critical multiplier for model perfor

  • Nvidia introduced NeMo Switchyard to route agent workflow steps across a model pool by quality, latency, and cost.

    NVIDIA (X) · attributed

    NVIDIA NeMo Switchyard helps developers route each agent workflow step across a chosen model pool based on their own quality, latency and cost criteria

  • Nvidia benchmarked 300+ verified skills and reported a 41% improvement in task correctness when skills were available.

    NVIDIA AI (X) · attributed

    NVIDIA conducted a controlled benchmark comparing agent performance with and without access to 300+ verified skills, finding a 41% improvement in task correctness

  • Zero Data Retention flags in coding agents do not prevent proprietary code from being transmitted to remote servers for inference.

    Arize AI Blog · attributed

    Zero Data Retention (ZDR) flags in coding agents are a compliance illusion that fails to prevent the initial transmission of proprietary code to remote servers.

Why it matters

For operators, Nvidia's AVO claim reframes capability gains as an infrastructure problem solvable with routing, skills, and orchestration rather than waiting solely for the next base-model generation.

Limits and uncertainties

The perfect ARC-AGI-3 score is a vendor-reported result on a public benchmark set, and the packet does not provide independent verification or production-task validation.

The Towards AI deconfliction article appears in the cluster with only a title and no excerpt or analysis in the packet.

Practical implications

Teams evaluating coding agents should weigh benchmark gains from scaffolding and skills against privacy limits such as ZDR not blocking initial code transmission to remote inference servers.

Multi-step agent builders may evaluate model routing tools like NeMo Switchyard and verified skill libraries as levers for cost, latency, and correctness before upgrading base models.

What to watch

Whether independent replication or third-party evaluation confirms AVO's reported 100% ARC-AGI-3 public-set result outside Nvidia's own reporting.

Whether Nvidia publishes fuller ARC-AGI-3 methodology details and whether verified-skills benchmarks translate to measurable gains on real internal codebases.

Sources

LLMgram editorial selection and synthesis · @llmgram. LLMgram is not the original publisher of this information.
Continue on LLMgram: Open in AI Signal →
Original reporting: Nvidia says its general-purpose coding agent system AVO scored 100% across all 25 environments in the ARC-AGI-3 public set, completing all 183 levels (Terry Chen/NVIDIA Technical…