Expert re-grading shows frontier models nearly ace physics benchmarks
How good are frontier models at physics?
Low scores on physics benchmarks suggest frontier language models struggle with advanced physics, but experts who audited six widely used benchmarks found most errors were not the models' fault. After correcting flawed reference solutions and ambiguous questions, GPT-5.6-Sol's mean@4 on HLE-Physics jumped from 47.3% to 78.7%, and its corrected pass@4 reached 94.4% on retained CritPt challenges. The findings indicate current benchmarks substantially understate model ability and are nearing saturation.
Most audited cases initially evaluated as incorrect reflect these benchmarking issues rather than errors in the models' physics reasoning.
- quantumtwist
The last author also wrote up a quite-readable blogpost here that accompanies the article: https://jsous.github.io/blogs/is-physics-dead/
One tidbit I found particularly interesting: "We tested a GPT-based agentic system, which previously succeeded in resolving several open mathematical conjectures, on open problems in theoretical physics. To our temporary relief and encouragement, it has not managed to fully resolve even one autonomously. We also noted that these agents made considerably less progress on the partially resolved physics problems than on open mathematics problems of comparable difficulty, both by our analysis and independent agent-based analysis of partial results in each domain."
- qt31415926
Article: "How Good Are Frontier Models at Physics?
Expert Re-Grading Reveals Broken Evaluations and
Near-Saturation of Leading Benchmarks"
John Sous from Yale posted a fairly solid study on how nearly all the physics benchmarks are broken, as they consistently mark correct answers as incorrect.
When hand grading instead, they found out that the models have actually already saturated the benchmarks which is a little bit scary.
- RomanKornev
Starting to feel more and more like chinese room experiment
The models are confidently answering physics questions, treating it as a math problem, but they don't fundamentally "get it" and even recently failed simple "should i drive to car wash" test
The sample efficiency is just crazy low
Still surprising that even with this they managed to saturate the benchmarks
- mch82
Are models able to do math now, or do they still rely on “tools” to do the math?
- amluto
(Trained physicist here)
From personal experience, frontier models absolutely struggle with understanding a physical situation based on words. (Okay, I haven't played with Astra much. GPT-5.6 Sol makes outrageous errors that anyone understanding a real world object would not make. And I was just asking it about NPT threads, not advanced physics.)
But seriously, what's up with these benchmarks? The example question in the paper is:
> PHYBench, problem 140: equivalent expressions for the same rope tension
> Problem statement. Three identical homogeneous balls are placed on a smooth horizontal surface, touching each other and are close enough to each other. A rope is wrapped around the spheres at the height of their centers, tying them together. A fourth identical sphere is placed on top of the three spheres. Find the tension T in the rope. It is given that the weight of each sphere is P.
For some reason the paper was focused on the fact that the grader didn't notice that some models were producing answers that were trivially algebraically equivalent to the reference answer. But this is missing the elephants in the room:
1. "touching each other and are close enough to each other": the right response is "hey, Professor, what do you mean 'close enough to each other'? They're sitting on a table in an equilateral triangle, all touching (i.e. tangent at their equators), right? Did you have a different configuration in mind?"
2. The answer is 0. Go find four baseballs or fours […]