ZeroHour
Story · 1 source · 1 articlefirst updated ()

Anthropic launches Claude Fable 5.1 and Mythos 5.1; OpenAI counters days later with GPT-6 Astra at matching $10/$50 pricing

infoModel releaseimportance 92
What's new: 2026-09-01: Anthropic launches Claude Fable 5.1 (with Mythos 5.1); Fable 5.1 more than doubles Fable 5's Terminal-Bench-Science 0.1 score (52.6% vs 24.7%)." }, { "date": "2026-09-01/02", "change": "Pricing and cost changes at launch: cache-read price cut 75% to $0.25 per million tokens and 1M-token context added, but 1.7x output token usage raises per-task cost ~20%." }, { "date": "2026-09-02",…
Merged summary · glm-5.3-flash · rewritten as coverage arrives

Anthropic released Claude Fable 5.1 and Mythos 5.1 as flagship coding and long-horizon work models, touting new benchmark highs (52.6% Terminal-Bench-Science 0.1, more than double Fable 5's 24.7%), a 1M-token context, and a 75% cache-price cut — though one…

Anthropic launched Claude Fable 5.1 alongside Mythos 5.1 (reported 2026-09-01/02), positioning Fable 5.1 for coding, knowledge work, and long-running problem-solving. Fable 5.1 scores 52.6% on Terminal-Bench-Science 0.1 versus 24.7% for Fable 5, 29.0% for Opus 5, and 22.4% for GPT-5.6 Sol; Latent Space separately reports 66 on the Artificial Analysis Intelligence Index (vs 63 for Opus 5), 59.1% on HLE, and 91.4% on Terminal-Bench v2.1 — distinct benchmarks from Terminal-Bench-Science 0.1. Pricing is $10/$50 per million input/output tokens with a 1M-token context and cache reads cut 75% to $0.25, but per-task cost rose ~20% because output token usage grew 1.7x, and ~4% of eval output tokens came from server-side fallback routing to Opus 4.8/Opus 5. The sources characterize the improvement differently: Latent Space calls it a new-SOTA launch with gains of +130 Elo on GDPval-AA v2, +122 Elo on AA-Briefcase, and +9 points on tau3-Banking over Fable 5, while Simon Willison's hands-on test (an animated pelican) notes other benchmarks improved only slightly, and Latent Space itself says Fable 5.1 is roughly tied with Opus 5 on some agentic knowledge-work measures. Release notes highlight Enterprise Frontier Safeguards and zero-data-retention support; community analysis (unconfirmed) suggested Fable and Mythos may share underlying weights with different safety/routing behavior. Two days later (reported 2026-09-03), OpenAI began rolling out GPT-6 Astra to a limited set of organizations, with planned availability for all ChatGPT Plus, Pro, Business and Enterprise users, the OpenAI API, and AWS. Astra's $10/$50 per million token API pricing matches Claude Fable 5 and 5.1, and OpenAI's self-reported benchmarks show it outperforming Fable on most measures, including 99.9% on ARC-AGI 3 (released in March) — though these are self-reported and not independently confirmed in the reports. Latent Space also noted the Fable launch coincided with a busy release window, with Grok 4.7 and Gemini Flash 3.8 still on the way.

  • Anthropic released Claude Fable 5.1 alongside Mythos 5.1, positioned as flagship models for coding, knowledge work, and long-running problem-solving (reported 2026-09-01/02).
  • Claude Fable 5.1 scores 52.6% on Terminal-Bench-Science 0.1, up from 24.7% for Fable 5, versus 29.0% for Opus 5 and 22.4% for GPT-5.6 Sol.
  • Latent Space reports Fable 5.1 at 66 on the Artificial Analysis Intelligence Index vs 63 for Claude Opus 5, 59.1% on HLE, and 91.4% on Terminal-Bench v2.1 (a different benchmark from Terminal-Bench-Science 0.1).
  • Latent Space-reported gains over Fable 5: +130 Elo on GDPval-AA v2, +122 Elo on AA-Briefcase, and +9 points on tau3-Banking.
  • Fable 5.1/Mythos 5.1 pricing: $10/$50 per million input/output tokens, 1M-token context window, cache reads cut 75% to $0.25.
  • Per-task cost rose ~20% due to 1.7x output token usage; ~4% of eval output tokens came from server-side fallback routing to Opus 4.8/Opus 5.
  • Release notes highlight Enterprise Frontier Safeguards and zero-data-retention support; community analysis (unconfirmed) suggested Fable and Mythos may share underlying weights with different safety/routing behavior.
  • Sources disagree on characterization: Willison describes gains on other benchmarks as slight rather than step-change, while Latent Space reports the launch as new SOTA; Latent Space also notes Fable 5.1 is roughly tied with Opus 5 on some…

Coverage timeline

  1. · 15d ago
    Simon Willison· 82
    Claude Fable 5.1 made me a really nice animated pelican

    Anthropic launched Claude Fable 5.1, claiming gains in coding and long-running tasks, with 52.6% on Terminal-Bench-Science 0.1.