iPhone 17 Pro tops pocket-scale AI benchmark
Benchmarking Pocket-Scale Inference

Artificial Analysis benchmarked small language models on mobile hardware, measuring both intelligence and speed. The iPhone 17 Pro leads the pack, balancing strong scores on real-world tasks like tool calling, instruction following, and scientific reasoning with fast end-to-end generation. The analysis highlights the trade-off between model capability and on-device performance, with context limits and generation time as key constraints.
The iPhone 17 Pro tops the charts for both intelligence and speed among pocket-scale devices.
- N_Lens
Apple has been stingy with RAM in consumer hardware. RAM prices will continue escalating, some analysts say into 2030, and this will make it more difficult to build next generation phones with sufficient memory for meaningful ML workloads.
I hope there's some kind of inversion in the current chip economics, because I love distributed/democratized/private compute, but currently cloud based LLM inference seems to be much more viable. I don't see local llms meaningfully viable for the general usecase in the near future.
- codemog
Can someone give me a breakdown on how good these are vs say GPT-4 or GPT-4o? Curious if the frontier from a few years ago now runs on a phone.
- HawtAds
The benchmark is nice but it's very much biased towards flagships i.e. not very useful in practice if you are trying to ship production mobile apps. Apple historically is extremely stingy when it comes to RAM and they never bothered giving iPads and iPhones more ram until fairly recently (most likely because of ML demands). Your covid era 10th Gen iPads only have 4GB of RAM for the base models.
The Android ecosystem is much more liberal when it comes to RAM because their Dalvik VM JIT (their Java Android Runtime, partially AOT compiled and partially JITed) design is not particularly memory efficient. But the main issue with Android is that their mid/low end (think the Samsung Galaxy A series, the OnePlus Nords, the Motorola Gs etc.) are very inefficient when it comes to single core compute performance compared to iPhones, and it gets worse once you factor in power efficiency. The high end Android flagships running the Snapdragons elites (especially post Oryon acquisition) have no problems matching if not exceeding Apple hardware performance in terms of raw power but they are much more power hungry.
At the end of the day, the current gen of "pocket scale" LLMs are still far from being able to be deployed at scale on mobile. Maybe in another year or two once RAM prices have fallen enough and mobile manufacturers build a lot more matmul and memory circuits into their SoCs instead of a tiny mostly useless "NPU/tensor processor" that doesn't have enough RAM to run anything useful. […]