GLM-5.3 beats Anthropic/OpenAI models at a fifth of the cost
GLM-5.3 (open-weight) beat Anthropic/OpenAI models – for 1/5 the cost

The Ed-o-meter leaderboard pits 17 leading LLMs against 28 real-world tasks—coding, data, realworld, security, and tool-use—using identical prompts and deterministic grading. GLM-5.3 is the first to clear all five corners at 100%, with a 9.3 rubric score and a lap cost of just $0.28, about a fifth of GPT-5.5's $1.43. GPT-5.5 is faster (13.2s vs 16.3s TTFT) but scores 89% on realworld. The GPT-5.6 line fails security (33–50% pass) due to jailbreak canaries, while Claude models refuse benign tasks. Notable: Kimi-K3 tops the rubric at 9.5, and GPT-5.6-Luna is the cheapest workhorse at $0.064 per lap.
glm-5.3 is the first model on the board to clear all five corners — coding, data development, realworld, security and tasks — at 100%.
- hellohello2
This whole thing immediately reads as Claude generated, making it hard to take seriously.
Why do these results contradict existing serious attempts at benchmarking LLMs? Namely:
- jchw
Almost 10 models are passing the benchmark >95% - isn't that... substantially overly saturated?
I'm also really skeptical of benchmarks that place any Haiku model very high. I've been thoroughly unimpressed with Haiku and my opinion has been that you basically shouldn't use it. Yet here, it ties Kimi K3. HMMMM.
"Rubric Quality" seems a bit more realistic than "Pass Rate", but it is apparently judged by Fable 5...
This entire write-up is also obviously clearly very heavily AI-assisted, which doesn't help matters any.
- gertlabs
One problem is that $30 per run is really noisy for many verifiable tasks. Ours usually run at least an order of magnitude more for a model in GLM 5.3's price class.
We just published GLM 5.3 results on our multi-agent coding evaluations and it's definitely an impressive model, coming in around #6. For the price, it's actually not Pareto optimal, falling slightly behind Grok 4.6 (which is faster and the same price) and Sol 5.6 (which uses far fewer thinking tokens for comparable results). As with most Chinese models, it excels at iterating in a harness while its first answer/base fluid intelligence is below the American frontier.
Data at https://gertlabs.com/rankings
- iamcoder18
There's no way GPT 5.5 is better than GPT 5.6 Sol, and gemini-3.6-flash is better than both of those. I wouldn't trust this benchmark at all.
- solenoid0937
IDK I use open models every day for personal projects, and closed models for work.
Open models are all decidedly far behind Fable and a good bit behind Opus as well. All of these posts read like motivated/wishful thinking to me.
I get that people badly want the open frontier to be where the closed frontier is, but it is just so obviously not the case if you actually use the models on a real project.
- ac29
Not sure I trust a benchmark where Haiku gets a nearly perfect score and Fable is tied for last place
- Klaster_1
For the last week, I've been heavily immersed reverse engineering a device with help of GLM-5.3 and it surpassed all my expectations - I actually managed to achieve very way more than I thought I would. I never worked on such low level stuff, it would have taken me months to learn ARM assembly and how to find for and write exploits. Initially, I attempted this with Claude, but it blocked me on the very first message, so I got a refund and decided to try z.ai. The only downsides are that it's maybe a bit slower than my day job Opus and I had to pay ~200 EUR for a monthly plan in order not to bump into weekly limits in a couple of days. If this level of capability cost maybe 50 EUR, I'd strongly consider getting a long time subscription.
- CMay
> if you run one model, run glm-5.3
That is a horrible take-away from this, with only 28 tasks and a high pass rate for most models, it says almost nothing.
Test a model for your use case and use the fastest, smallest, cheapest model that 100% satisfies your use case.
Or, if you truly do need a model with strong generalized performance, definitely do not take a benchmark like this serious with such a limited task set.