Opus 5.5 may be getting quietly worse — this benchmark is watching
Livenerf: Has Opus 5.5 been nerfed yet?
livenerf is a 30-day, pre-registered benchmark that tracks whether Claude Opus 5.5 degrades after its 2026-09-22 launch. Using a frozen panel of 78 questions that the model only sometimes answers correctly, it runs daily on a Claude Max subscription via headless Claude Code, logging everything to detect statistical drift. Early validation shows that lower effort settings reduce output tokens far more than accuracy, and that a same-family model swap (Opus 5) is not distinguishable from Opus 5.5 in a single validation run.
The secondary signal I care most about is the output token count per sample. If a model quietly starts thinking less, this is where it shows up first, often before accuracy moves at all.
- jug
We also have Nerf Bench:
https://www.bridgebench.ai/nerf-bench
They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
- giancarlostoro
I think Anthropic tries to adjust things to handle their user load, which has negative effects. For example, before they added 1 million tokens back in February, I could effectively prompt Claude to keep going until x number of requirements were completed, now I have to make a loop, not only that, but Claude will finish before my loop time sometimes and wait for the next pass, which could be in like 30 minutes or so. This probably helps them keep a lighter load, but it could probably tick off anybody, the output is the same, you just aren't getting it as quickly as you once did. I don't hate it, I use Claude on my off-hours to work on personal side projects.
- msejas
As a claude code power user, when I get the 'rate the feedback on Claude' pop up, I used to say good or fine out of habit, and immediately after sending this feedback, I felt an instant degradation and mistakes that usually don't happen.
Now I dismiss it every time and the quality is more consistent.
Complete adhoc and personal experience but something I've observed, wouldn't be surprised if they nerfed on a per session basis
- lxgr
Glad to see good old human cognitive biases (or, in very advanced cases, just sloppy methodology) are still highly competitive with SOTA model hallucinations.
- sheepscreek
> It could also mean nothing happened and people are pattern-matching on noise.
It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.
The compounding effect can definitely cause temporary regressions in some domain or use-case that doesn’t have really good coverage in their internal tests.
What makes this particularly challenging in the case of LLMs is how changing some language in the prompt can vastly affect the output.
But this is less true as models become larger and additional parameters are able to capture each and every possible nuance of the language. I’d say caching and cache tuning or token optimization/tweaks to thinking are the biggest culprits today.
- johnfn
"Nerf"ing models isn't real in the vast majority of reported cases. Benchmarks like this or the 100 other "let's see if nerfing is real" copies would have shown it by now if it was.
I made a graphic to explain why people feel like the models get nerfed:
https://x.com/thesilenceturns/status/2103551351825543610
The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there's a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.
- khalic
Given the tens of bot accounts in this thread making obviously misleading statements, I think you’re onto something here. Might be a good idea to open a Patreon or something, you probably need throwaway accounts to prevent anthropic from feeding you good models.
- rw2
I 100% believe the models are being nerfed and feel it when I use it. Launch day LLM + the week it's released is fantastic. After they get all the press the nerfing starts. When they release a new model sometimes it's not significantly better, it's just not nerfed.
- bandrami
Is it the models or the users' dopamine receptors that get nerfed after a few weeks?
- semiquaver
One thing that these nerfing conspiracy theorists fail to realize is that Anthropic isn’t the only organization that directly serves Claude inference.
My company uses Claude models exclusively via Azure and AWS bedrock, which have their own licensed copies of the weights.
If all these people are so convinced Anthropic is nerfing models, have they tried other inference providers? Do they think the nerfing is coordinated across independent providers? Why wouldn’t any of these nerf-benches use these comparison points or even talk about them?
I think the most likely conclusion by far is that this is a psychological phenomenon.