Local AI can now handle 88.7% of real-world queries, study finds
Intelligence per Watt: Measuring Intelligence Efficiency of Local AI
Researchers evaluated 20+ local language models on 1M real-world queries, measuring accuracy, energy, latency, and power. They introduce intelligence per watt (IPW) as a unified metric. Local models answered 88.7% of queries, and IPW improved 5.3x from 2023 to 2025. Local accelerators achieved at least 1.4x better IPW than cloud accelerators on identical models, suggesting local inference can meaningfully redistribute demand from centralized infrastructure.
Local accelerators achieve at least 1.4x lower IPW than cloud accelerators running identical models, revealing significant headroom for local accelerator optimization.
- jmiskovic
Incredibly important research. We've reached the point where local LLMs are good enough! It takes less time for local model to take the first action on your task than it does for Claude to validate your login, put you into queue and start issuing the commands. Local models are persistent and 100% predictable unlike any cloud offering. It's better for the power system for the demand to be distributed. During the winter time the GPU also doubles as a 300W in-house heater. Not to mention avoiding personal data collection and re-selling.
- WASDx
I tried calculating historical "intelligence per cost" recently but stopped when I realized intelligence is not linear. For any meaningful "x per y" you can just double "y" if you have a half as efficient system to get the same result but so-called intelligence doesn't work like that.
- surprisetalk
I rather like this paper, but I think it is generous to say their benchmarks measure "intelligence".
[0] https://arxiv.org/pdf/1911.01547
We have no dang clue what intelligence is, nor how to measure it.
- scottcha
Very cool paper and some interesting things for local serving. Though its a little apples-to-oranges we do publish live energy stats for all models on our service here https://portal.neuralwatt.com/energy-pricing in case you are interested in what this looks like on the cloud side. FWIW DSV4.1 flash is really getting popular due to its IPW.
Some of the items like model routing, if you do it per request instead of per session, can break down on the cloud from an energy and cost POV since one of the best things you can do for both is to maintain the KV cache which both reduces time component of energy and the quite expensive prefill energy.
I am keen on the future where we have local/cloud hybrid serving which is cache aware. I do think that could be the best use of energy resources for AI.
- polotics
Did I miss something or does this article not bother to indicate how much RAM their M4-Max had?