NVIDIA AVO scores 100% on ARC-AGI-3 interactive reasoning benchmark
NVIDIA reported that its general-purpose coding agent, NVIDIA AVO, achieved a perfect score on the ARC-AGI-3 interactive reasoning benchmark, clearing all 183 levels across 25 public environments without instructions, explicit rules, or stated goals. The company framed AVO as an agent that inspects, plans, implements, and evaluates while using memory, tools, and execution feedback to accumulate learning during play. If the ARC-AGI-3 result holds under wider scrutiny, it would signal unusually strong interactive generalization for a coding-oriented system rather than a narrow task-specific model. The claim currently rests on NVIDIA's own post on X, without independent replication details or peer-reviewed publication in the available packet, so builders should treat the headline result as promising but self-reported until corroborated.
NVIDIA AVO scores 100% on ARC-AGI-3 interactive reasoning benchmark
NVIDIA says its general-purpose coding agent scored 100% on the ARC-AGI-3 interactive reasoning benchmark. NVIDIA AVO completed all 183 levels across all 25 public environments with no instructions, explicit rules, or stated goals.
Key takeaway
NVIDIA AVO's reported perfect ARC-AGI-3 run suggests coding agents can solve novel interactive puzzles without pre-specified goals or rules.
What happened
NVIDIA said on X that NVIDIA AVO, described as a general-purpose coding agent, scored 100% on the ARC-AGI-3 interactive reasoning benchmark.
According to the post, AVO finished all 183 levels across 25 public environments with no instructions, explicit rules, or stated goals, using inspect-plan-implement-evaluate loops with memory, tools, and execution feedback.
Evidence
NVIDIA AVO scored 100% on the ARC-AGI-3 interactive reasoning benchmark.
X · attributed
NVIDIA says its general-purpose coding agent scored 100% on the ARC-AGI-3 interactive reasoning benchmark.
AVO completed all 183 levels across all 25 public environments without instructions, explicit rules, or stated goals.
X · attributed
NVIDIA AVO completed all 183 levels across all 25 public environments with no instructions, explicit rules, or stated goals.
NVIDIA AVO continuously inspects, plans, implements, and evaluates using memory, tools, and execution feedback.
X · attributed
NVIDIA AVO continuously inspects, plans, implements, and evaluates, using memory, tools, and execution feedback to build on what it learns along the way.
Why it matters
Self-reported benchmark milestones from model vendors rarely settle capability debates until methods, costs, and independent scores are published.
Limits and uncertainties
The available reporting is limited to NVIDIA's X post, with no independent verification cited in the packet.
Some source snippets in the packet are truncated, leaving incomplete wording around how AVO figured out tasks.
Practical implications
Teams benchmarking agentic coding systems should compare ARC-AGI-3 methodology, environment coverage, and failure modes before treating vendor scores as production readiness.
Operators evaluating autonomous agents should weigh inspect-plan-implement-evaluate loops with tool use against their own safety and observability requirements.
What to watch
Independent ARC-AGI-3 scores or technical write-ups from parties other than NVIDIA.
Publication of full run logs, compute costs, and per-environment failure or retry details for AVO.
Original reporting: Our general-purpose coding agent just scored 100% on the ARC-AGI-3 interactive reasoning benchmark.
NVIDIA AVO completed all 183 levels across all 25 public environments, figuring