Anthropic appears to be A/B testing reduced effort levels in Claude Code

Anthropic appears to be A/B testing reduced effort levels in Claude Code

A developer reports that Anthropic is running a server-side A/B test on Claude Code versions 2.1.236 and later, shrinking the effort scale so that 'high' now maps to 10 out of 100—the same value 'low' used to have. Older versions and Opus 5 are unaffected. The changelog doesn't mention this change, leaving users confused about degraded performance.

Since 2.1.237 the model reads 'high' effort as 10 out of 100, the exact number 'low' used to be and the changelog doesn't say a word.
  1. pizzafeelsright

    Whatever Opus 5 is doing should not happen.

    Prompt was "read and update the config file with new data". This work on 4.6 takes <2

    minutes to read the file, parse the new data, and patch.

    Opus 5 Result: 43 minutes of pulling containers, running sandboxes, creating testing suites, which included evaluating the entire repo beyond the scope of the config file.

    Both: one file modification

  2. trq_

    Hi all, Thariq from the Claude Code team here. I posted this on Twitter, but just reposting here:

    We sometimes test API serving configs in Claude Code before rolling them out, and one running now maps the numerical effort value differently.

    That's why Claude may tell some of you it's at "10" on high. The scale isn't 0-100, the number isn't meaningful on its own, and the effort you selected is the effort you're getting. We've run in-depth evals to confirm this doesn't affect model performance.

    This should be the same experience, but if you see a clear regression please hit /feedback and send me the ID. Will give credits.

  3. boredumb

    Not specifically Anthropic but why are we allowing billing to take place in tokens that are nebulous and fully controlled by the operators who have no aligned incentives?

    If I have a user input and then sanitize and inject that into a prompt to do something, I have no idea how much that is going to cost at all and no real way to measure this properly. A parallel example is digital ocean or aws, i can go and measure/limit my compute/fs/memory/startup times/etc and while it can be impossible to get down to the last flop of money allocated - i can run things on a real budget with real constraints, opposed to an LLM where I have to .. prerun a sanitized user prompt through a tokenizer and then ask an LLM to guess what it may do and give token consumption estimates and then act on those in any sane manner for the user?

    Perhaps i'm missing something to do realistic and static rails on things but I don't see a serious way at scale to use the token billing model handling things requiring a users free text input short of having to go pander to VC money to throw money at it until someone else figures it out.

    *to clarify my rambling...

    We should be billed and given controls based on resource usage itself and not an opaque token concept on top of not being able to spin any knobs that control it's resource usage.

  4. hpone91

    Update from Thariq on twitter. https://x.com/trq212/status/2091247114869432543

    "We sometimes test API serving configs in Claude Code before rolling them out, and one running now maps the numerical effort value differently.

    That's why Claude may tell some of you it's at "10" on high. The scale isn't 0-100, the number isn't meaningful on its own, and the effort you selected is the effort you're getting. We've run in-depth evals to confirm this doesn't affect model performance.

    This should be the same experience, but if you see a clear regression please hit /feedback and send me the ID. Will give credits."

  5. monideas

    This phenomenon was so bad and so noticeable with Fable that I downgraded my Max subscription ($200) to pro ($20). It’s basically useless. Codex 5.6 Sol is actually very good, I’ll just create another account to get more usage

  6. Insimwytim

    LLM users don't want to put in effort, so they offload tasks to LLM.

    LLM doesn't seem to be keen to put in effort either!

    Is this AGI?

  7. N_Lens

    I suspect it's not just this, there's plenty of 'optimization' around rubberbanding usage limits as well as routing to a different model in the backend. The incentives are too strong.

  8. ricardobeat

    I use Opus 5 almost exclusively at low effort, and get good results. Especially on high it seems to go out on completely unasked-for tangents. Same seems to be true for Sonnet 5. Older models did not behave like this.

    The mood change in just six months is wild, in February this year Claude was the most liked LLM by far.

More from this day

2026-08-22