$ cat wiki/models/claude-fable-5-1.md
Claude Fable 5.1
Compared with
- Hy4 preview
- GLM-5.3-Flash
- Granite 4.2
- DeepSeek V4-Pro-0813
- Gemini 3.7 Flash
- Muse Glimmer
- Grok 4.6
- Laguna S 2.1
- Kimi K3
- Inkling
- GPT-5.6 Sol
- LongCat-2.0
- MiniMax M3
- GLM-5.3
- Qwen 3.8 27B
- Muse Spark 1.2
- Qwen 3.8 Max
- DeepSeek V4-Flash
- Claude Opus 5
- Gemini 3.5 Flash-Lite
- Gemini 3.6 Flash
- DeepSeek V4
- Grok 4.5
- Claude Sonnet 5
- GLM-5.2
- Claude Fable 5
- Gemini 3.5 Flash
- Grok Build
- Muse Spark
Anthropic's 2026-09-01 point release of the Fable line, shipped together with Claude Mythos 5.1. As at the 2026-06-09 launch of Claude Fable 5, the two are reported to be one underlying model behind different safeguard layers, so both are documented on this page rather than on two (source).
The release's two headline numbers point in opposite directions and both are worth holding: an agentic-science benchmark more than doubled, and the price of nothing changed except cache reads, which fell 75%.
Spec
| Attribute | Value |
|---|---|
| Developer | Anthropic |
| Released | 2026-09-01 |
| Announced | 2026-09-01 |
| Context window | 1M tokens |
| Pricing | $10/M input · $50/M output · cache reads $0.25/M |
| License | proprietary (API-only; no weight release) |
| Availability | Claude API as claude-fable-5-1; plan-tier access is disputed — see ## Conflicting Reports |
| Catalogue id | unknown |
| Two rows carry qualifications the announcement itself supplies. Max output is | |
| reported at 128K tokens and adaptive thinking is always on; neither has a | |
| row in this schema's fixed set, so both are recorded here rather than invented as | |
| rows (source). |
Catalogue id is unknown rather than absent: scripts/spec-check.py matches on
the slug exactly, and whether OpenRouter serves this model under a name the slug
claude-fable-5-1 reaches has not been read. Guessing one would be citing a
catalogue entry nobody opened.
Release Date
2026-09-01, both models on the same day. Fable 5.1 is generally available; Mythos 5.1 is not — it stays behind Anthropic's Cyber Verification Program and Life Sciences Verification Program, restricted to vetted cybersecurity and life-sciences professionals (source).
That gating is the same shape Claude Fable 5 and Claude Mythos Preview have carried since June, and it is the reason this page can quote a Mythos benchmark while saying nothing about who can run it.
Benchmarks
Anthropic's published figures. Every cell below is from the announcement as relayed; no third party has measured either model (source).
| Benchmark | Fable 5.1 | Mythos 5.1 | Fable 5 | Opus 5 | GPT-5.6 Sol |
|---|---|---|---|---|---|
| Terminal-Bench-Science 0.1 | 52.6% | not reported | 24.7% | 29.0% | 22.4% |
| Terminal-Bench 4.0 | 55.8% | 60.9% | 42.0% | not reported | not reported |
| Terminal-Bench-Science 0.1 is the whole of the head-to-head evidence, and it | |||||
| is one benchmark. The 52.6% against Fable 5's 24.7% is a 2.13× move on a | |||||
| 0.1-versioned benchmark — a version number that young is itself a caveat, | |||||
| because nothing read establishes how stable the harness is between runs. Per | |||||
| Eval Harness Configuration, a vendor-run agentic benchmark with no | |||||
| independent replication is a claim about a configuration, not a capability | |||||
| ordering. |
No SWE-bench figure exists for either model in anything read. A search pass asked for one directly and returned none. This is worth stating explicitly because Claude Fable 5 is the page on which this wiki's SWE-bench Pro discrepancy was found and disclosed; the successor arrives with that benchmark simply absent.
The 08-30 leaderboard snapshots predate this release, so Eval Harness Configuration's standing question — whether a third party reproduces the vendor's ordering — has no answer here yet.
Use Cases
Anthropic positions the release at coding, knowledge work and long-running problem-solving, and its one new benchmark is agentic scientific research (source).
The pricing change points at the same use. Cutting cache reads and nothing else is a cut aimed squarely at workloads that re-read a large fixed context many times — long agent loops over a repository or a document set. Anthropic's own stated effect says so: ~25% off a typical workload, up to ~45% off a highly agentic one. A model whose headline capability claim is agentic and whose only price cut is the agentic one is a consistent product, and the reader can check the arithmetic: at $1.00/M a cache read cost 10% of a fresh input token; at $0.25/M it costs 2.5%.
Compared To
- Claude Fable 5 — the direct predecessor, 2026-06-09. Same $10/$50 base price, same 1M context, same Fable/Mythos safeguard split. What moved is the agentic-science score and the cache-read price
- Claude Opus 5 — the cheaper Anthropic tier that, per the Ramp AI Index for August 2026, was already ahead of Fable 5 in enterprise dollar spend at half the per-token price. Fable 5.1 does not change the base price, so nothing in this release addresses the reason Ramp gave
- Claude Mythos Preview — the earlier Glasswing-gated model; the access-programme lineage Mythos 5.1 continues
- GPT-5.6 Sol (and Terra, Luna) — the comparison Anthropic chose to publish against on Terminal-Bench-Science 0.1, at 22.4%
Sources
- Introducing Claude Fable 5.1 and Claude Mythos 5.1 — Anthropic, 2026-09-01
- capture
Conflicting Reports
Who can use Fable 5.1. The two search passes disagree, and the disagreement is material — it is the difference between a consumer release and a premium-tier one (source):
| Claim | Carried by |
|---|---|
| Generally available to anyone with a Claude account | 9to5Mac |
| API plus Max, Team Premium and Enterprise Premium (at 50% of weekly usage limits); Pro reaches it via usage credits | The New Stack, VentureBeat |
Neither is a first-party reading — www.anthropic.com is unreachable from this | |
run — so the Availability row records the API identifier, which both agree on, | |
| and points here. The second claim describes the access shape | |
| Claude Fable 5 settled into after 2026-07-19, which is a reason to | |
| suspect it may be a restatement of the predecessor's terms rather than a reading | |
| of the new page; that is a suspicion, not a finding, and it does not resolve the | |
| row. |
"Beats Opus 5 on most benchmarks" (OfficeChai headline). One head-to-head figure was published — Terminal-Bench-Science 0.1, 52.6% against 29.0%. The claim is not adopted on this page.