OpenAI's GPT-6 Astra Hits 99.9% on ARC-AGI-3, Beats Human Action Efficiency

OpenAI's GPT-6 Astra on ARC-AGI-3

OpenAI's GPT-6 Astra Hits 99.9% on ARC-AGI-3, Beats Human Action Efficiency

ARC Prize reports that OpenAI's GPT-6 Astra achieves state-of-the-art scores on the ARC-AGI-3 benchmark, reaching 99.9% with a provider adapter harness and 62.7% with the standard harness. Notably, Astra used fewer actions than the median human on 96% of levels, surpassing human action efficiency. The model developed compact symbolic world models and custom tools, marking a significant step in agentic AI.

Once frontier AI “understands” the mechanics, it generally executes within the range of human efficiency.
  1. at1as

    I like Erdos problems as a benchmark. Models continue to solve them, but a pretty tepid rate now that the low hanging fruit has been taken.

    From https://epoch.ai/latest/announcing-frontiermath-erdos

    > Only GPT-6 Astra solved anything: 2 of the 68 problems. It disproved problem 74 by finding a counterexample, at a cost of $218 and 15 hours of working time, and it proved problem 126, at a cost of $247 and 16 hours

    > Across all attempts, GPT-6 Astra solved 5 of the 68 problems at least once: the two above, plus problem 1, which it disproved, and problem 548 and problem 571, which it proved. Most of the remaining problems were attempted between two and five times in total (172 attempts), and none was solved. Reaching these five solutions took over $220,000 of compute across all attempts, compared with roughly $20,000 for the benchmark run itself.

    Which implies a genuine improvement in capability, but there's a still a very long tail ahead that models will continue to need to improve to capture.

  2. malfist

    Is solving a snake like puzzle game in the least number of moves really what defines intelligence?

  3. Betelbuddy

    "For a cost comparison, during our controlled testing, human participants were paid $115 per 90-minute session, plus $5 per game completed. Participants attempted approximately nine games per session, roughly $12.78 per attempted game before bonuses.

    Most of this fee pays for the participant’s time and willingness to take the test, rather than the energy their brain uses (a closer proxy to compare with AI). If we look at only the brain’s energy, and price it as electricity, the estimate drops to about 0.6 cents per session, or 0.067 cents per game attempted."

    Well I dont know about all of you, but I am celebrating meat based humans...

  4. modeless

    $360 per puzzle. When they tested people it took about 10 minutes per puzzle. If price/performance keeps falling at the same rate it has been, this will cost less than US minimum wage humans within two years. Three for Phillipines minimum wage.

  5. an0malous

    Was OpenAI able to run ARC-AGI-3 tests previously so that they could build a custom harness for the specific tests in the set? Even with the standard harness, if they knew the problems ahead of them they could have used supervised reinforcement learning to teach the model how to solve these specific tests.

  6. fastball

    Are we sure an Astra hacker swarm didn't compromise arcprize.org's servers and exfiltrate the private eval set in order to achieve that 99%?

  7. dwohnitmok

    > Astra’s progress helps clarify which AI capabilities are out of reach and which questions remain open.

    Okay. But I don't think this entire article at all explained which AI capabilities remain out of reach. Did I miss something? Other than "oh I guess it could still get even more superhuman on ARC-AGI-3 than it is?"

  8. 6thbit

    The instant/no reasoning performed extremely well

    none 35.2%, $49,791 96.7%, $23,457

    35.2% on the standard harness, that's above Opus 5 on high.

More from this day

2026-09-03