GLM-5.3-Flash: A New Benchmark in AI Intelligence and Cost Efficiency
GLM-5.3-Flash Intelligence, Performance and Price Analysis
Artificial Analysis evaluates GLM-5.3-Flash, a new model from Zhipu AI, against the Intelligence Index v4.1.1, which combines nine benchmarks including GDPval-AA v2, Terminal-Bench v2.1, and Humanity's Last Exam. The analysis highlights the model's competitive intelligence score, cost per task, and token efficiency, positioning it as a strong contender in the AI landscape.
The Artificial Analysis Intelligence Index v4.1.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity's Last Exam, GPQA Diamond, CritPt, AA-Omniscience, AA-LCR.
- smartbit
Looking at these numbers IMHO, with Gemini you get the speed what you pay for:
Intel Cost
lig per
ence Task Speed
Gemini 3.7 flash 56 0.40 338
GLM 5.3 flash 57 0.09 49
Factor 1 4.4 6.9
- CGamesPlay
Several factual errors about the model here. The input modalities are listed as text only, but the headline feature is image support. The context length should be 1048576 (so should GLM-5.3's, also wrong on the charts).
- yipinwong
The analysis is still not compelling for me to switch
from gtp5.6-luna to GLM-5.3-flash given
- costs per task $0.05 vs $0.09
- speed 130 vs 88
- where GLM has only 5 more intelligence point: at this point few point is meaningless for most of models
https://artificialanalysis.ai/models/comparisons/glm-5-3-fla...
Been using Luna exclusively since the price drop, and i've been very satified with all tasks from planning, writing code, and other agent tasks. (just change thinking level from low <-> ultra)
---
btw, I did try out Ox Alpha, the coding feels good but still not way better for me to switch to it.
- m_ke
So how exactly is Anthropic and OpenAI ever going to pay back the trillions that they plan on spending?
- AnodicElegy
Impressive. It kicked everything between itself and Sol xhigh out of the Pareto frontier. Can't wait to try it out.