ZeroHour
Latent Spacepublished ()ingested 2

[AINews] Claude Fable/Mythos 5.1: new SOTA model, 75% cache price cut but 70% more output tokens

infoModel releaseimportance 92
AI summary · glm-5.3-flash

Anthropic launched Claude Fable 5.1 and Mythos 5.1, claiming new SOTA benchmarks, with 75% cache-read price cut and 1M-token context.

Anthropic released Claude Fable 5.1 and Mythos 5.1 as flagship models for coding and knowledge work, with a 1M-token context window and pricing of $10/$50 per million input/output tokens and cache reads cut 75% to $0.25. Artificial Analysis Intelligence Index scored Fable 5.1 at 66 versus 63 for Claude Opus 5, with HLE at 59.1% and Terminal-Bench v2.1 at 91.4%, though per-task cost rose ~20% due to 1.7x output token usage. Community analysis suggested Fable and Mythos may share underlying weights with different safety/routing behavior, and release notes highlighted Enterprise Frontier Safeguards and zero-data-retention support.

  • Fable 5.1 positioned for autonomous long-horizon work; Mythos 5.1 paired for knowledge work
  • Benchmark gains: GDPval-AA v2 +130 Elo, AA-Briefcase +122 Elo, tau3-Banking +9 points over Fable 5
  • On agentic knowledge work Fable 5.1 is roughly tied with Opus 5 on some measures
  • ~4% of eval output tokens came from server-side fallback routing to Opus 4.8/Opus 5
  • Launch coincides with competing releases: OpenAI Astra, Grok 4.7 and Gemini Flash 3.8 on the way
Full article2,972 words · extracted from latent.space · click to collapse

With Astra clearly finally warming up for a full launch (with @sama and @openai writing about it again after a month of self imposed pacing ), there’s a familiar window to take the narrative with the round robin of model launches, with Grok 4.7 and Gemini Flash 3.8 also on the way. But that’s also perhaps not the best way to frame today’s launch… which got well over 12M views updating the sitting world best model yet again:

The benchmark table speaks for itself:

While per-token pricing is the same as Fable/Mythos 5, the cache reads had a 75% price cut … great news for long sessions/long context users, however offset by observed 1.7x output token usage increases per Artificial Analysis, for a total net per-task cost increase of 20% (see recap below).

Also don’t World Labs’ Astra launch , by far the most impressive world model launch we’ve ever seen, and on a regular day would have easily gotten title story cards. You can catch up on Fei Fei and Justin Johnson’s vision on our pod and trace from Marble to Astra and what we were talking about with the true potential of world models:

AI News for 8/31/2026-9/1/2026. We checked 12 subreddits, 544 Twitters and no further Discords. AINews’ website lets you search all past issues. As a reminder, AINews is now a section of Latent Space . You can opt in/out of email frequencies!

AI Twitter Recap

Top Story: Fable 5.1 and Mythos 5.1 release and reactions

What happened

Anthropic launched Claude Fable 5.1 and Claude Mythos 5.1 as its new flagship models for coding and knowledge work.

Anthropic announced the release directly, positioning them as “the world’s most advanced models for coding and knowledge work” via @claudeai

Anthropic product/engineering voices framed Fable 5.1 specifically around autonomous, multi-step work: “complex, multi-step work that runs on its own,” with emphasis on coding, knowledge work, and long-running problem solving via @mikeyk

Anthropic kept list pricing for Fable 5.1 at $10 / $50 / $12.5 per million tokens for input / output / cache write, while cutting cache read price by 75% to $0.25 / MTok , again noted by @mikeyk , @Teknium , and independently quantified by @ArtificialAnlys

Early benchmark screenshots and system-card excerpts drove much of the discussion, especially around Terminal-Bench-Science, SWE-family evals, HLE, FrontierCode, and Artificial Analysis via @StevenDillmann , @scaling01 , @ArtificialAnlys

A key interpretive claim emerged from community analysis: Fable and Mythos 5.1 may be the same underlying weights, with different safety/routing behavior , not different base models, per @eliebakouch and later @nrehiew_

User reactions split along multiple axes: very strong praise for coding/planning ability and tone, but complaints around rate limits, safeguards false positives, subscription UX, and unclear benchmark presentation via @danshipper , @theo , @kimmonismus , @GregKamradt , @kylebrussell , and @eliebakouch

Official claims and model positioning

Anthropic’s own messaging was straightforward: Fable 5.1 is for difficult, delegated, long-horizon work, while Mythos 5.1 is the paired release for knowledge work. The main official launch post is @claudeai . Supporting commentary from Anthropic staff emphasized:

autonomous long-running tasks via @mikeyk

improved honesty / better failure reporting (“when it’s stuck it says so instead of reporting success”) via @mikeyk

new enterprise-oriented controls, especially Enterprise Frontier Safeguards (EFS) , positioned as “ZDR++” for agent observability in enterprise environments via @alexalbert__

zero-data-retention support highlighted by users as an important adoption unlock, especially @danshipper

The official pitch was not merely “better benchmark model,” but “usable autonomous worker” — fast enough, cheap enough in cached agent settings, and enterprise-compatible enough to deploy.

That positioning mattered because Fable 5 had a reputation — repeated in reactions — for being powerful but sometimes impractical. Dan Shipper summarized the prior criticism as Anthropic having “built a supergenius in a datacenter that was almost unusable,” then argued 5.1 addresses slowness, verbosity, and awkward tone via @danshipper .

Technical details and numbers

Core published/priced details

From @ArtificialAnlys :

Context window: 1 million tokens

Modalities: text + image inputs

Pricing: unchanged from Fable 5 for

input: $10 / 1M tokens

output: $50 / 1M tokens

cache write: $12.5 / 1M tokens

Cache read price: reduced from $1.00 to $0.25 / 1M tokens ( 75% cut )

Artificial Analysis notes this cache cut materially benefits agentic workloads where much of the prompt is repeatedly re-read from cache.

Artificial Analysis headline results

Also from @ArtificialAnlys :

Artificial Analysis Intelligence Index: 66 at max effort

ahead of:

Claude Opus 5 max: 63

Claude Fable 5 max: 62

GPT-5.6 Sol max: 61

Grok 4.6 high: 61

HLE: 59.1%

previous best cited: Fable 5 at 55.5%

Terminal-Bench v2.1: 91.4%

SciCode: 62.0%

τ³-Banking: +9 points over Fable 5

GDPval-AA v2: 1853 Elo , +130 over Fable 5

AA-Briefcase: 1694 Elo , +122 over Fable 5

But AA also adds an important qualification:

On agentic knowledge work, Fable 5.1 is effectively tied with Opus 5 on some measures, not obviously dominant

Their eval used Anthropic’s default server-side fallback , with safety-flagged requests routed to Claude Opus 4.8 or Claude Opus 5

Fallback accounted for ~4% of output tokens across the Intelligence Index

That fallback detail became one of the most consequential technical caveats in community interpretation.

Cost per task

Artificial Analysis also reported:

Fable 5.1 max: $3.76/task

Fable 5 max: lower, so 5.1 is 20% more expensive per task

reason: Fable 5.1 uses ~1.7× output tokens

cache cut saves ~$1.40 per task

Fable 5.1 xhigh: score 65 , cost $2.72/task

Opus 5 max: score 63 , cost $2.34/task

This produced one of the key tensions in the reaction cycle: Fable 5.1 looks clearly better at the frontier ceiling, but not clearly better on every cost-efficiency framing.

Additional framing from @nicdunz :

Fable 5.1 Max: 66 intelligence , 140M tokens , $3.69/task

Fable 5 Max: 62 , 83M tokens , $3.14/task

GPT-5.6 Sol Max: 61 , 70M tokens , $0.95/task

This post argues Sol remains the clear winner on intelligence-per-dollar and intelligence-per-token, even if Fable 5.1 wins absolute ceiling.

Benchmark snippets from system-card discussion

Community members extracted several benchmark points:

From @StevenDillmann :

Terminal-Bench-Science 0.1

Fable 5: 24.7%

Fable 5.1: 52.6%

more than 2× improvement

From @scaling01 :

DeepSWE: 67.4%

FrontierCode 1.1 Extended: 63.6%

FrontierSWE v2: 0.57 , “highest of the models Proximal evaluated”

From @Sauers_ :

Humanity’s Last Exam: 65% with tools

From @perplexity_ai :

Perplexity’s August WANDR evaluation:

score 0.601

$12.76 per task

21% higher score

37% lower cost than Fable 5

From @scaling01 :

Artificial Analysis Intelligence Index score 66 , “back on the frontier”

From @theo :

cache price cut was the “biggest W”

in CursorBench , costs were cut by “almost 50% ” while scoring higher

From @kimmonismus :

Fable 5.1 High appears stronger and cheaper than Sol 5.6 Max on Cursor Bench

though this is a secondary paraphrase, not an original benchmark report

From @scaling01 :

Mythos 5.1 displays verbalized grader awareness in 65% of long agentic coding environments

That last point is especially interesting: it suggests the model may explicitly model the evaluator in a large fraction of long-horizon coding contexts, which raises both capability and eval-gaming questions.

Safeguards and routing details

Two tweets capture the technical interpretive crux:

@eliebakouch : “Fable and Mythos 5.1 are the EXACT same weights” , with internal activations used for safety classification and escalation to a bigger classifier, then fallback to Opus 4.8 for dangerous requests

@nrehiew_ : if true, the difference is “likely the threshold set for the safeguard classifier”

These are not official Anthropic statements in the tweet corpus, but they line up with the official AA note that fallback routing served ~4% of output tokens on AA’s evals via @ArtificialAnlys .

This led to repeated community questions about whether benchmark lines reported as “Mythos” versus “Fable” are genuinely comparable, especially if one naming convention mostly indicates which safety path was active , not which base model was doing the work. See @eliebakouch , @eliebakouch , and @eliebakouch .

Facts vs opinions

Facts strongly supported by official/independent sources

Anthropic launched Claude Fable 5.1 and Claude Mythos 5.1 via @claudeai

Fable 5.1 pricing retained $10 / $50 / $12.5 for input/output/cache write, with cache reads cut to $0.25 / MTok via @mikeyk and @ArtificialAnlys

Fable 5.1 has 1M context , image+text input support, and tops AA’s Intelligence Index at 66 via @ArtificialAnlys

AA’s evaluation included server-side fallback , with ~4% of output tokens served by fallback models via @ArtificialAnlys

Fable 5.1 showed very large gains on several coding/agentic benchmarks, including 52.6% on Terminal-Bench-Science via @StevenDillmann

Plausible but not fully verified claims

Fable and Mythos 5.1 are identical weights with different safeguard/routing behavior via @eliebakouch and @nrehiew_

Some benchmark labels may reflect safety mode / route differences rather than separate base-model performance via @eliebakouch

“It talks like a normal person now” / reduced “Claudese” is widely reported anecdotally, but is still subjective, despite some lexical stats below

Opinions / subjective judgments

“Strongest coding model we’ve used” from @danshipper

“Fable is the frontier model by a good margin right now” from @AravSrinivas

“Astra is going to absolutely destroy Fable 5.1” from @scaling01

“I honestly haven’t noticed much difference compared to Fable 5” from @kimmonismus

Text extracted automatically; images, tables and formatting may be missing. Original: https://www.latent.space/p/ainews-claude-fablemythos-51-new