NVIDIA AVO scores 100% on ARC-AGI-3 interactive reasoning benchmark

Nvidia AVO scores 100% on the ARC-AGI-3 interactive reasoning benchmark

NVIDIA AVO scores 100% on ARC-AGI-3 interactive reasoning benchmark

NVIDIA's general-purpose coding agent AVO scored 100% on the ARC-AGI-3 interactive reasoning benchmark, completing all 183 levels across 25 public environments without instructions, explicit rules, or stated goals. The company says AVO continuously inspects, plans, implements, and evaluates, using memory and tools to sustain progress on long-running tasks.

Our general-purpose coding agent just scored 100% on the ARC-AGI-3 interactive reasoning benchmark.
  1. mellosouls

    Underlying article should be the link:

    https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-...

    There you will find the extremely important qualifier it's the public set, not the private set (with the risk of overfitting, ie the results not repeating when submitted to be run in competition), and the detail that this is essentially a harness added to Opus 5, not Nvidia's own models.

    Obviously still impressive, you would think.

  2. woeirua

    I thought ARC-AGI-3 was explicitly a test of raw model performance excluding the harness? Adding the harness back in doesn't tell us anything new. We've known that agents are capable of long horizon reasoning with sufficient harnesses. GPT-4(?) was capable of beating Pokemon 18 months ago but models only became capable of beating it without a harness in the last six months...?

  3. antinucleon

    AVO’s paper author (ex-NVIDIAN) is here. This work was done half a year ago for GPU kernels, and the same approach has now been applied to ARC-AGI-3. I think people are still underestimating the evolution progress; e.g., recently we made a self-improving evolution harness that generated an entire inference stack and is better than SGLang/vLLM on various tasks: https://int21.ai/insights/addressing-the-inference-bottlenec...

  4. subzel0

    The 100% score was achieved on the 25 public set, not on the semi-private or private sets.

  5. magicalhippo

    The blog post: https://developer.nvidia.com/blog/nvidia-avo-reaches-100-on-...

    Using Claude Opus 5, but it can use others:

    AVO is also designed to operate across frontier models. While our full public-set result used Claude Opus 5, we additionally paired AVO with GPT-5.6 Sol on a challenging subset of games. In these limited experiments, Sol reached matched levels faster in wall-clock time in several cases, while Opus used fewer environment actions in matched-level comparisons. These preliminary results suggest complementary operating profiles across models, and we leave a broader systematic comparison to future work

More from this day

2026-08-21