Claude Fable 5.1: Longer agentic runs, cheaper cache reads, and stricter thinking-block rules
Claude Fable 5.1 and Claude Mythos 5.1

Anthropic's Claude Fable 5.1 succeeds Claude Fable 5 at the same price, with cache reads at a quarter of the cost. It targets long-running agentic coding, multistep research, and document work, and brings a 1M-token context window, always-on adaptive thinking, and per-message effort control. Breaking changes: forced tool use is unsupported, earlier models can't read its thinking blocks, and editing earlier turns invalidates them. New betas include turn-scoped system messages and progress updates between tool calls. Text output is watermarked, and images/videos carry C2PA credentials.
To keep thinking valid across a long session, treat the conversation as append-only.
HN discussion
1331- felixrieseberg
(I work at Anthropic)
Beyond all the benchmarks, I think Fable 5.1 is a big improvement in writing style. It sounds a lot less stereotypically like other Claude models, has (imho) a much more natural style, and responds to my style instructions more reliably. More work to be done (and we will!) but reading better prose makes me so much happier.
Another point I expect not to get much attention until it all happens at once is science. People have been correctly excited about the many "sudden" breakthroughs LLMs are making in Maths, but some of the science benchmarks make me believe we'll soon see similar developments in other scientific domains. Fable 5.1 more than doubled Fable 5's Terminal-Bench-Science [1] score, which I think is meaningful.
[1] https://github.com/harbor-framework/terminal-bench-science
- simonw
Pelicans for thinking effort low, medium, high and xhigh (that xhigh one is pretty good): https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
I'm still waiting for effort max to finish.
EDIT: I fixed a bug in my tooling so it now records summarized reasoning traces - here's that max pelican, which is a significant improvement: https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
Took just under 14 minutes to generate, and at 65927 output tokens cost me a hefty $3.30!
Excerpts from the reasoning trace:
> Adding pedal shapes near both feet, with the far foot on the second leg partially visible behind the frame. I'm considering whether to add a small scarf or cap for extra character, but leaning toward keeping it simple to avoid clutter.
> Now I'm debating a bicycle helmet on the head versus the pelican's signature crest—the beak and pouch already read clearly as "pelican," so a helmet could reinforce the bicycle theme without losing identity, though it might compete with the crest for visual space.
> I realize the beak at (484,84) would overlap with the dome helmet, so I need to shrink the helmet so it only covers the top of the head, adjusting its arc endpoints to sit higher and narrower so the beak can attach cleanly at the front without collision. [...]
> I'm adding a darker tip region to represent the primary feathers, then reconsidering the trailing edge to include scalloped feather curves instead of one smooth line for a more natural look. [...]
> N […]
- GodelNumbering
The price reduction comes from the cache read pricing falling from $1/M to $0.25/M, which means that Fable 5.1 now costs half of Opus's cache read costs ($0.5/M).
This gives a lot of credit to the theory that Anthropic did not get much bite on Fable at its original pricing, which in turn likely places a ceiling on LLM pricing in general.
Interestingly also, if you take away terminal-Bench-Science 0.1 results, it is hard to see ANY improvement:
Terminal-Bench 4.0: Fable 5.1 is +3.5% vs Opus 5.
GDPval-AA v2: +1.5% vs Opus 5.
OSWorld 2.0: +2.5% vs Opus 5.
Humanity's Last Exam (with tools): +1.6%
Keep in mind that this is supposed to be an entirely higher tier of a model than Opus 5. For one tier up and one version up, these are not really improvements. Probably leaves no room to place Opus 5.1 anywhere. Combined with the fact that they are selling 'readability'... Has frontier progress finally stalled?
- mlaux
Looks like all three breaking changes are patches for inadvertent chain of thought disclosure. Someone found out (don't have the tweet handy) that if you created a bogus "think_deeply" tool and then forced the model to use it, it would output what is believed to be its raw thinking there - I believe the first breaking change stops this. The second two are aimed at people getting Haiku to repeat thinking blocks from other models verbatim (since it can see the decrypted version). I get that in their eyes it's an "exploit" but still kinda disappointing that they patched this
- exabrial
Anyone ever seen the SouthPark episode making fun of Game of Thrones: A Song of Ass and Fire? Anthropic's announcements reminds me of "The Dragons Are Coming" running joke.
What they have done:
* Nerfed Fable, as many of noted it's useless
* Leverage Mythos as a marketing strategy, claiming its too good to release
* Removed thought traces, one of the only useful things to make sure your prompts are working correctly
* Continue tons of hype about how good they are without delivering, going to great lengths to publish how their model "hacked" its way out of a sandbox they misconfigured.
* Push a bunch of EU Overregulation onto the rest of the world with text watermarking, decreasing quality of answers
Last year, they were at least focused on making improvements. Nowadays its just a bunch of handwaving at the church of how good they are.
The only saving grace is Opus 4.6 is still available. Just sucks we haven't seen any measurable improvement, despite all of the ceremony.
- madrox
I am finding that I am now less interested in better models than I am in token budgets. My issue with Anthropic models now is that I don't feel like I can rely on them as a daily driver because they'll dry up before my quota resets.
I am becoming dependent on AI to make a living, and I need predictable spend on it. If I know I can't use a model regularly all month, my enthusiasm is limited.
I urge Anthropic to get better at this aspect of their business so I can come back to it.
- tarr11
“ Claude Fable 5.1's writing is generally a step up from earlier Claude models, with fewer stock phrases and less unexplained jargon. In some cases, though, its prose is denser than Claude Fable 5's: sentences run longer and there are fewer paragraph breaks.”
I cancelled my pro max Claude subscription last week; codex is much more succinct. I am curious if this is getting better.
I don’t think Anthropic realizes that humans have a token limit too and it can be exhausting to read Claude’s output. Prose density is not the same thing as succinctness.
- pookieinc
"Price. Fable 5.1 will cost an estimated 25% less than Fable 5 for typical workloads, wherever usage is billed by token. This is because we’re reducing our pricing on cache reads (where the model reads inputs that have already been processed and stored). For highly agentic work, the savings will often be much larger—up to approximately 45%."
Glad to see this!
- kccqzy
I’ll be very excited to try it out and see the actual improvement in writing style. The denser writing style probably won’t bother me.
Anthropic seems to be listening to community complaint on HN about how the writing style is grating. And apparently the solution from Anthropic is to add this block to every conversation!?
> Mannered prose substitutes metaphor and flourish for direct statement. Instead of "a parameter worth varying," the mannered writer produces "a dial worth turning." Instead of "this point still matters," they write "this point earns its keep." The phrases exist to display the writer, not to convey the idea, and readers can tell. That is why mannered prose irritates: it makes the reader work harder so the writer can perform. It is also imprecise. Metaphors drag in connotations the writer did not choose and cannot control. The fix is to say what you mean. When a literal phrase is available, use it.
The above was quoted verbatim from https://platform.claude.com/docs/en/build-with-claude/prompt...
- jumploops
> For example, in testing by the investment firm Millennium, Fable 5.1 found the cause of a rare crash on their internal systems that none of their engineers (or any other model) had been able to explain after several years of trying.
Say what you will about LLM-generated code, but stories like this give me hope that software will never be as buggy as it once was.