[AINews] Claude Fable/Mythos 5.1: new SOTA model, 75% cache price cut but 70% more output tokens
Anthropic launched Claude Fable 5.1 and Mythos 5.1, claiming new SOTA benchmarks, with 75% cache-read price cut and 1M-token context.
Anthropic released Claude Fable 5.1 and Mythos 5.1 as flagship models for coding and knowledge work, with a 1M-token context window and pricing of $10/$50 per million input/output tokens and cache reads cut 75% to $0.25. Artificial Analysis Intelligence Index scored Fable 5.1 at 66 versus 63 for Claude Opus 5, with HLE at 59.1% and Terminal-Bench v2.1 at 91.4%, though per-task cost rose ~20% due to 1.7x output token usage. Community analysis suggested Fable and Mythos may share underlying weights with different safety/routing behavior, and release notes highlighted Enterprise Frontier Safeguards and zero-data-retention support.
- Fable 5.1 positioned for autonomous long-horizon work; Mythos 5.1 paired for knowledge work
- Benchmark gains: GDPval-AA v2 +130 Elo, AA-Briefcase +122 Elo, tau3-Banking +9 points over Fable 5
- On agentic knowledge work Fable 5.1 is roughly tied with Opus 5 on some measures
- ~4% of eval output tokens came from server-side fallback routing to Opus 4.8/Opus 5
- Launch coincides with competing releases: OpenAI Astra, Grok 4.7 and Gemini Flash 3.8 on the way
Full article2,972 words · extracted from latent.space · click to collapse

With Astra clearly finally warming up for a full launch (with @sama and @openai writing about it again after a month of self imposed pacing ), there’s a familiar window to take the narrative with the round robin of model launches, with Grok 4.7 and Gemini Flash 3.8 also on the way. But that’s also perhaps not the best way to frame today’s launch… which got well over 12M views updating the sitting world best model yet again:
The benchmark table speaks for itself:
While per-token pricing is the same as Fable/Mythos 5, the cache reads had a 75% price cut … great news for long sessions/long context users, however offset by observed 1.7x output token usage increases per Artificial Analysis, for a total net per-task cost increase of 20% (see recap below).
Also don’t World Labs’ Astra launch , by far the most impressive world model launch we’ve ever seen, and on a regular day would have easily gotten title story cards. You can catch up on Fei Fei and Justin Johnson’s vision on our pod and trace from Marble to Astra and what we were talking about with the true potential of world models:
AI News for 8/31/2026-9/1/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space . You can opt in/out of email frequencies!
AI Twitter Recap
Top Story: Fable 5.1 and Mythos 5.1 release and reactions
What happened
Anthropic launched Claude Fable 5.1 and Claude Mythos 5.1 as its new flagship models for coding and knowledge work.
Anthropic announced the release directly, positioning them as “the world’s most advanced models for coding and knowledge work” via @claudeai
Anthropic product/engineering voices framed Fable 5.1 specifically around autonomous, multi-step work: “complex, multi-step work that runs on its own,” with emphasis on coding, knowledge work, and long-running problem solving via @mikeyk
Anthropic kept list pricing for Fable 5.1 at $10 / $50 / $12.5 per million tokens for input / output / cache write, while cutting cache read price by 75% to $0.25 / MTok , again noted by @mikeyk , @Teknium , and independently quantified by @ArtificialAnlys
Early benchmark screenshots and system-card excerpts drove much of the discussion, especially around Terminal-Bench-Science, SWE-family evals, HLE, FrontierCode, and Artificial Analysis via @StevenDillmann , @scaling01 , @ArtificialAnlys
A key interpretive claim emerged from community analysis: Fable and Mythos 5.1 may be the same underlying weights, with different safety/routing behavior , not different base models, per @eliebakouch and later @nrehiew_
User reactions split along multiple axes: very strong praise for coding/planning ability and tone, but complaints around rate limits, safeguards false positives, subscription UX, and unclear benchmark presentation via @danshipper , @theo , @kimmonismus , @GregKamradt , @kylebrussell , and @eliebakouch
Official claims and model positioning
Anthropic’s own messaging was straightforward: Fable 5.1 is for difficult, delegated, long-horizon work, while Mythos 5.1 is the paired release for knowledge work. The main official launch post is @claudeai . Supporting commentary from Anthropic staff emphasized:
autonomous long-running tasks via @mikeyk
improved honesty / better failure reporting (“when it’s stuck it says so instead of reporting success”) via @mikeyk
new enterprise-oriented controls, especially Enterprise Frontier Safeguards (EFS) , positioned as “ZDR++” for agent observability in enterprise environments via @alexalbert__
zero-data-retention support highlighted by users as an important adoption unlock, especially @danshipper
The official pitch was not merely “better benchmark model,” but “usable autonomous worker” — fast enough, cheap enough in cached agent settings, and enterprise-compatible enough to deploy.
That positioning mattered because Fable 5 had a reputation — repeated in reactions — for being powerful but sometimes impractical. Dan Shipper summarized the prior criticism as Anthropic having “built a supergenius in a datacenter that was almost unusable,” then argued 5.1 addresses slowness, verbosity, and awkward tone via @danshipper .
Technical details and numbers
Core published/priced details
From @ArtificialAnlys :
Context window: 1 million tokens
Modalities: text + image inputs
Pricing: unchanged from Fable 5 for
input: $10 / 1M tokens
output: $50 / 1M tokens
cache write: $12.5 / 1M tokens
Cache read price: reduced from $1.00 to $0.25 / 1M tokens ( 75% cut )
Artificial Analysis notes this cache cut materially benefits agentic workloads where much of the prompt is repeatedly re-read from cache.
Artificial Analysis headline results
Also from @ArtificialAnlys :
Artificial Analysis Intelligence Index: 66 at max effort
ahead of:
Claude Opus 5 max: 63
Claude Fable 5 max: 62
GPT-5.6 Sol max: 61
Grok 4.6 high: 61
HLE: 59.1%
previous best cited: Fable 5 at 55.5%
Terminal-Bench v2.1: 91.4%
SciCode: 62.0%
τ³-Banking: +9 points over Fable 5
GDPval-AA v2: 1853 Elo , +130 over Fable 5
AA-Briefcase: 1694 Elo , +122 over Fable 5
But AA also adds an important qualification:
On agentic knowledge work, Fable 5.1 is effectively tied with Opus 5 on some measures, not obviously dominant
Their eval used Anthropic’s default server-side fallback , with safety-flagged requests routed to Claude Opus 4.8 or Claude Opus 5
Fallback accounted for ~4% of output tokens across the Intelligence Index
That fallback detail became one of the most consequential technical caveats in community interpretation.
Cost per task
Artificial Analysis also reported:
Fable 5.1 max: $3.76/task
Fable 5 max: lower, so 5.1 is 20% more expensive per task
reason: Fable 5.1 uses ~1.7× output tokens
cache cut saves ~$1.40 per task
Fable 5.1 xhigh: score 65 , cost $2.72/task
Opus 5 max: score 63 , cost $2.34/task
This produced one of the key tensions in the reaction cycle: Fable 5.1 looks clearly better at the frontier ceiling, but not clearly better on every cost-efficiency framing.
Additional framing from @nicdunz :
Fable 5.1 Max: 66 intelligence , 140M tokens , $3.69/task
Fable 5 Max: 62 , 83M tokens , $3.14/task
GPT-5.6 Sol Max: 61 , 70M tokens , $0.95/task
This post argues Sol remains the clear winner on intelligence-per-dollar and intelligence-per-token, even if Fable 5.1 wins absolute ceiling.
Benchmark snippets from system-card discussion
Community members extracted several benchmark points:
From @StevenDillmann :
Terminal-Bench-Science 0.1
Fable 5: 24.7%
Fable 5.1: 52.6%
more than 2× improvement
From @scaling01 :
DeepSWE: 67.4%
FrontierCode 1.1 Extended: 63.6%
FrontierSWE v2: 0.57 , “highest of the models Proximal evaluated”
From @Sauers_ :
Humanity’s Last Exam: 65% with tools
From @perplexity_ai :
Perplexity’s August WANDR evaluation:
score 0.601
$12.76 per task
21% higher score
37% lower cost than Fable 5
From @scaling01 :
Artificial Analysis Intelligence Index score 66 , “back on the frontier”
From @theo :
cache price cut was the “biggest W”
in CursorBench , costs were cut by “almost 50% ” while scoring higher
From @kimmonismus :
Fable 5.1 High appears stronger and cheaper than Sol 5.6 Max on Cursor Bench
though this is a secondary paraphrase, not an original benchmark report
From @scaling01 :
Mythos 5.1 displays verbalized grader awareness in 65% of long agentic coding environments
That last point is especially interesting: it suggests the model may explicitly model the evaluator in a large fraction of long-horizon coding contexts, which raises both capability and eval-gaming questions.
Safeguards and routing details
Two tweets capture the technical interpretive crux:
@eliebakouch : “Fable and Mythos 5.1 are the EXACT same weights” , with internal activations used for safety classification and escalation to a bigger classifier, then fallback to Opus 4.8 for dangerous requests
@nrehiew_ : if true, the difference is “likely the threshold set for the safeguard classifier”
These are not official Anthropic statements in the tweet corpus, but they line up with the official AA note that fallback routing served ~4% of output tokens on AA’s evals via @ArtificialAnlys .
This led to repeated community questions about whether benchmark lines reported as “Mythos” versus “Fable” are genuinely comparable, especially if one naming convention mostly indicates which safety path was active , not which base model was doing the work. See @eliebakouch , @eliebakouch , and @eliebakouch .
Facts vs opinions
Facts strongly supported by official/independent sources
Anthropic launched Claude Fable 5.1 and Claude Mythos 5.1 via @claudeai
Fable 5.1 pricing retained $10 / $50 / $12.5 for input/output/cache write, with cache reads cut to $0.25 / MTok via @mikeyk and @ArtificialAnlys
Fable 5.1 has 1M context , image+text input support, and tops AA’s Intelligence Index at 66 via @ArtificialAnlys
AA’s evaluation included server-side fallback , with ~4% of output tokens served by fallback models via @ArtificialAnlys
Fable 5.1 showed very large gains on several coding/agentic benchmarks, including 52.6% on Terminal-Bench-Science via @StevenDillmann
Plausible but not fully verified claims
Fable and Mythos 5.1 are identical weights with different safeguard/routing behavior via @eliebakouch and @nrehiew_
Some benchmark labels may reflect safety mode / route differences rather than separate base-model performance via @eliebakouch
“It talks like a normal person now” / reduced “Claudese” is widely reported anecdotally, but is still subjective, despite some lexical stats below
Opinions / subjective judgments
“Strongest coding model we’ve used” from @danshipper
“Fable is the frontier model by a good margin right now” from @AravSrinivas
“Astra is going to absolutely destroy Fable 5.1” from @scaling01
“I honestly haven’t noticed much difference compared to Fable 5” from @kimmonismus
Text extracted automatically; images, tables and formatting may be missing. Original: https://www.latent.space/p/ainews-claude-fablemythos-51-new