Fable 5 crushes nanoGPT speedrun, closing 81.7% of the gap to human record

NanoGPT Speedrun Frontier

Fable 5 crushes nanoGPT speedrun, closing 81.7% of the gap to human record

Prime Intellect ran 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun. Fable 5, using Claude Code, achieved a record of 2,726, closing 81.7% of the gap to the human baseline of 2,600. The leaderboard shows Opus 5, Kimi K3, and GPT-5.6 Sol following, with detailed traces and equal-budget comparisons available.

We ran 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun.
  1. lhl

    Pretty interesting to see on the training front. I've used most of these models to grind semi-autonomously (days at a time) on kernel optimizations (except for Fable - it kept triggering guardrails almost immediately and bouncing me down to Opus 4.8 at the time). I think for a lot of people that might be the biggest problem, although it looks like Opus 5 still does well.

    I found that if you leave them alone undirected, the models (especially GPT models) will rathole, but with the right scaffolding it seems to work pretty well. My general loop is to start with ideation and profiling phase, limit # of runs before forcing moving on to the next item down the list, and then iterating, potentially mixing models with "fresh eyes". This is probably something that could be fully automated, but I like checking in once a day or so and seeing what's happening and redirecting.

  2. vibe42

    "Almost every model finds the same winning ideas. What separates the best traces is what an experiment leaves behind. They preserve weak signals long enough to validate them, but they also have a better understanding of the results."

    Curious if a harness that helped preserve signals in some history log would change the outcome.

    Also curious if different goal prompts would have changed the outcome. Not a bunch of prompt engineering; small diffs like "consider novel solutions, keep track of weak signals".

    IMO they allocated quite a bit of GPU time to the same goal prompt.

  3. nsingh2

    What's going on with sol here? The note says it spends a lot of time waiting, did it just not effectively use time (i.e. something like parallel runs) so it's graph ends up being stretched in the time axis?

    I'm also seeing notes like on Opus 5 saying it was a run with a older serial version of program.md, so the graphs aren't complete apples-to-apples comparisons?

    Edit: the blog seems to address these https://www.primeintellect.ai/blog/measuring-autonomous-rese...

  4. totetsu

    “We ran 153 autonomous runs across 18 frontier models on the nanoGPT optimizer speedrun.”

    Uh.. okay.. but whats a run… read blog

    “We want to measure how well frontier models can conduct research….””we ran 153 autonomous runs on the nanoGPT optimizer speedrun across”

    Okay but what is a optimiser run and what connection does it have to being good at research?

    “For comparison, Anthropic's internal automated AI R&D evaluation optimizes a model on a CPU node,”

    So I should go look what Anthropic was doing to understand?

    Why not just explain what it means in their blog..

  5. espadrine

    I would be interested to have a third X-axis with dollars.

    Time is sometimes more about inference infrastructure (especially with systolic chips) than model quality (and providers tweak knobs to support higher batches at the expense of latency).

    Tokens are not always fully equivalent between models.

More from this day

2026-08-23