Recreating Minecraft Is Not a Benchmark

Recreating Minecraft Is Not a Benchmark

The author argues that viral demos like recreating Minecraft in one prompt have become 'demo-benchmarks'—easy to overfit and measure preparation rather than capability. They point to Thinking Machines' Inkling Small scoring within a point of its flagship on the Artificial Analysis Intelligence Index with a third of the parameters, showing that public static tests leak into training data. They suggest holdout evals like LiveBench or ARC-AGI as better alternatives, but acknowledge demos' appeal for their instant clarity.

A test you can perfect on a schedule measures preparation instead of capability, to me that’s anti the very definition of a benchmark, it should be a hard test, something very hard to perfect.
  1. bluegatty

    Right answer wrong thesis.

    The example of 'labs are optimizing for this' is the wrong thesis.

    AI arbitrarily generating something of 'apparent sophistication' is not that hard - being able to produce it to spec that has invariable vague elements - and then being able to rationally modify it is the problem.

    Look at image gen: you press the magic button and get a 'Pelican on a Bike' - but you can't just change the 'hat' of the Pelican. You have to press the magic button again, and you get a whole different Pelican on a Bike.

    This is the fundamental conceit.

    It's akin to the conceit that 'writing the code is the work' - but it's not - it's the research, the design, docs, integration and all that 'know-how' that has to 'sit somewhere' so the shape of the thing output can be adapted and moulded.

    Arbitrary code output is not quite worthless but almost, it's the the '80% that means another 80% and then anther 80% to go'. It's like a nicer stating point.

    The research capabilities of the AI, which don't make for nice demos, are arguably more powerful.

  2. 0xb0565e486

    I keep seeing Astra make beautiful 3d stuff online, yet when I feed it some old school RuneScape assets (even tried with some very detailed guidelines) and asked it to generate some new plausible assets it failed horribly.

    I think there’s still something really off with current (frontier) models when it comes to creating “novel” stuff? Even 2004 style graphics..

    Or am promoting it wrong?

  3. orbital-decay

    This sounds unconvincing, because a) pelican test is subjective, there's simply nothing to leak as it has no available direct answers and maybe an extremely faint preference signal, and b) the same small models actually do perform well when you change the subject. Some models are genuinely trained to be better at some domain, such as 2D layouts or vector graphics in this case. It all depends on particular recipes and datasets. Which is the actual reason these tests are poor as vibe checks: they don't do anything to disentangle generalization, memorization, and training preference. One-shotting popular software in particular is definitely not a good test of anything as memorization is going to dominate it.

    AAII is also not very useful, neither is any generic score/benchmark. If you want a weather forecast you aren't looking at the average temperature of Earth.

    (actually when did the term "one-shot" get hijacked to mean something other than "one example"?..)

  4. hombre_fatal

    > That’s the problem, these tests can’t tell you how good a model is anymore because it’s trivial for labs to optimise for exactly these tests by the next release.

    But they don't prove the claim. Are the models amazing at recreating Minecraft, but the second you swap the word Minecraft out with another game or a custom game, it shits the bed?

    That's not what I see. My feed is full of people using Astra to recreate all sorts of games from Diablo to some random idea they came up with, in ridiculously polished detail like animations that would have taken me weeks of iteration in gpt-5.6-sol but it was a single shot by Astra.

  5. mrkramer

    >GPT-6-Astra-Max

    "Generate a SVG of a PlayStation 4 controller as nicely done as you can"

    Time taken: 11m 54s

    No image was used

    I was expecting 10s not 12m; if it takes 12m to do it, I would rather try to do it myself. And what it means "No image was used", wasn't it trained on bunch of images including probably console controllers images?

More from this day

2026-09-06