Intelligence costs dropped 56x in six months — what a 100x drop changes

What Happens When the Cost of Intelligence Drops 100x

Intelligence costs dropped 56x in six months — what a 100x drop changes

The cost of a given level of AI intelligence has fallen 56x in under six months, from $1.22 per task in February to $0.022 today, according to data from Artificial Analysis. The author, who runs a data analysis project, found that a full pass over 10,000 papers now costs about $100, down from several thousand dollars in March. This price collapse is turning previously infeasible analyses into viable projects, and the rate of decline is accelerating.

The ceiling unlocks new kinds of tasks; the floor unlocks volume.
  1. Balooga

    Jevons Paradox [1]

    > when technological improvements that increase the efficiency of a resource's use lead to a rise, rather than a fall, in total consumption of that resource.

    [1] - https://en.wikipedia.org/wiki/Jevons_paradox

    Las Vegas replaced the expensive incandescent lighting on the strip with cheaper to run LED equivalents. But the costs didn't come down because they were able to add more lights and larger displays.

    I think the same will happen with tokens. As the cost of tokens comes down, these models will just consume more tokens.

  2. andai

    > Reading everything becomes the default. At a cent per document, a model can read every paper

    I love how in our day "reading everything" means "the computer reads it for me".

    I expect soon the computer will be able to go on bicycle rides, and spend time with my wife.

  3. jbotdev

    I think speed is actually going to be a bigger factor than cost. Even projects where “money is no object” often hit a wall with LLM response times.

    Sure you can speed things up with parallel work under subagents, but as with parallelizing traditional computational tasks, there are diminishing gains.

    I keep hearing people saying just change the way you work to trust long-running agents and multi-task more, because they’re too slow to work with interactively for many use cases. I think that’s painful in a world where we expect humans to still heavily guide and interact with agents for their day-to-day work.

  4. AnotherGoodName

    I think a big one is robotics. A robot can today fold your laundry. It takes ~10mins per item. Seriously. It takes a long time to process the image find the corner move the claw to the corner of the shirt and attempt to straighten before folding.

    Robots right now generally move at glacial speeds. You might have seen robots doing flips in semi controlled environments but watch how slowly they open doors etc. processing time is a major bottleneck.

  5. newAccount2025

    I’m loving small models. The gemma4 26/31b models have been deeply impressive on weird prose analysis tasks that I am working on. Nova-micro is really stupid but is extremely fast when it’s smart enough to do something. I’m trying to be disciplined about able to evaluate quality vs cost everywhere for real systems built on this stuff. I probably need to get off Bedrock because it’s missing a lot of other little models that might be good competitors.

  6. MichaelNolan

    100x seems like an underestimate. Even with no model improvements, we should see that sort of reduction. Looking at TSMC’s margins, Nvidia’s margins, and OAI/Anth (alleged) margins on inference, there is a room for a 100x reduction.

    Right now all three of those are at abnormally high levels. Competition will come for all three.

  7. vanuatu

    I think what a lot of people miss about jevon's paradox is the elasticity of demand of the underlying resource

    textiles had jevons paradox, and many more textile workers were employed even when textile machines were being created, until we saturated the demand for cheap clothing in the world and then textile workers were kaput (same for farming, and horses)

    software is currently undergoing jevons paradox, but it's very unknown how high the ceiling of demand for software is. web dev might be doomed, but software in general i think is probably limitless

    Intelligence is also probably unbounded (atm software and intelligence are very closely tied together). its very possible token spend rides up the curve forever.

  8. bdhdhduuyd

    Personally I still see LLMs as very advanced search engines which lack intelligence. To me it seems that the cost of getting data is reduced by LLMs, not the cost of intelligence. I mean: we tell the model what we want to achieve, and the model responds with the right data in de form of code in seconds.

    That's why 'stackoverflow programmers' will have a hard time competing with LLMs but engineers are still needed for their intelligence.

    Well that's just my 2 cents.

More from this day

2026-08-21