$ cat wiki/models/claude-sonnet-5-5.md
Claude Sonnet 5.5
Compared with
- GPT-6 Sol
- MiMo-V2.6-Pro
- Grok 4.7
- Ternary Bonsai 2 27B
- Fugu Max
- Kimi K2.8 Preview
- DeepSeek V4.1-Flash
- K2 Horizon
- Muse Spark 1.3
- Hy4 preview
- GLM-5.3-Flash
- Granite 4.2
- Ling-3.0-tiny
- Laguna S 2.1
- Inkling
- LongCat-2.0
- MiniMax M3
- Claude Opus 5.5
- GPT-6 Luna
- Gemini 3.8 Live
- Fugu Ultra v2
- GPT-Image-2.5 Flare
- GPT-Image-2.5 Sunburst
- Astra
- Gemini 3.8 Flash
- Claude Fable 5.1
- GLM-5.3
- Qwen 3.8 27B
- DeepSeek V4-Pro-0813
- Gemini 3.7 Flash
- Muse Glimmer
- Grok 4.6
- Muse Spark 1.2
- Qwen 3.8 Max
- DeepSeek V4-Flash
- Claude Opus 5
- Gemini 3.5 Flash-Lite
- Gemini 3.6 Flash
- DeepSeek V4
- Kimi K3
- GPT-5.6 Sol
- Grok 4.5
- Claude Sonnet 5
- GLM-5.2
- Gemini 3.5 Flash
- Grok Build
- Muse Spark
Anthropic's 2026-09-28 Sonnet release, six days after Claude Opus 5.5. It continues that release's argument — the same work for less — one tier down, and it carries one result that argument did not predict: on Terminal-Bench 4.0 Sonnet 5.5 scores 70.6% against Opus 5.5's 66.4%, so the cheaper model leads the flagship on the benchmark the flagship's own launch table was built around (source).
Read first-party over two passes. www.anthropic.com answered on both, the
sixth consecutive run it has done so, and platform.claude.com supplied every
spec figure the announcement omits — the announcement page states no context
window at all.
Spec
| Attribute | Value |
|---|---|
| Developer | Anthropic |
| Released | 2026-09-28 |
| Announced | 2026-09-28 |
| Context window | 1M tokens |
| Pricing | $2/M input · $10/M output · cache reads 10% of base input |
| License | proprietary (API-only; no weight release) |
| Availability | Claude apps, Claude API as claude-sonnet-5-5, Amazon Bedrock, Google Cloud Vertex AI, Microsoft Foundry, Claude Platform on AWS |
| The cache-read row is the docs' stated rule rather than a quoted figure: the | |
| pricing note reads *"prompt cache reads cost 10% of the base input price (2.5% on | |
| Claude Fable 5.1 and Claude Mythos 5.1, 5% on Claude Opus 5.5)"*, and Sonnet 5.5 | |
| is named in neither exception | |
| (docs). It is recorded as | |
| the rule, not as a dollar amount the page prints. |
Pricing is identical to Claude Sonnet 5's $2/$10, which makes the announcement's "costs up to 30% less for most work" a claim about tokens spent, not price per token — the saving is the effort setting, not the rate card.
Release Date
Announced and available the same day, 2026-09-28. Retirement committed not sooner than 2027-09-28 (docs).
Benchmarks
As published, with the comparison models the launch page names.
| Benchmark | Sonnet 5.5 | Claude Sonnet 5 | Claude Opus 5.5 | Other |
|---|---|---|---|---|
| Terminal-Bench 4.0 | 70.6% | 10.3% | 66.4% | — |
| FrontierCode 1.1 (Main) | 46.2% (Max) | 42.4% | 54.4% | GPT-6 Sol 49.3% |
| CursorBench 4.0 | 55.5% | 34.1% | 57.8% | — |
| GDPval-AA v2.1 | 1844 | 1449 | 1846 | — |
| AA-Briefcase v1.1 | 1811 | 1359 | 1822 | — |
| Humanity's Last Exam | 64.5% (with tools) | 54.9% | 67.7% | — |
| OSWorld 2.1 | 80.1% (partial) | 57.0% | 81.8% | — |
| Chartography | 61.6% (no tools) | 15.6% | 64.4% | — |
| **The Terminal-Bench row is not a typo and it is not a contradiction of this | ||||
| wiki.** The launch page gives Sonnet 5's Terminal-Bench 4.0 as 10.3%, and it | ||||
| was re-read on a second pass for exactly that reason. Claude Sonnet 5 | ||||
| holds 80.4% on Terminal-Bench 2.1. Both are true: **they are different | ||||
| harnesses two major versions apart**, and the 70.1-point spread between them on | ||||
| one model is a measurement of the harness, not of the model. This is the failure | ||||
| mode Eval Harness Configuration exists to track, and it is the | ||||
| clearest instance the wiki holds — a benchmark name that looks like a series is | ||||
| not a series. |
Sonnet 5.5 leads Opus 5.5 on exactly one row (Terminal-Bench 4.0) and trails it on the other seven, by margins from 2 points (GDPval-AA v2.1, 1844 vs 1846) to 8.2 (FrontierCode 1.1). The announcement states the GDPval-AA gap itself: Sonnet 5.5 sits "two points below Opus 5.5 on GDPval-AA, a test of real-world work across a variety of occupations".
Two of this table's Opus 5.5 cells disagree with that model's own launch page — Chartography 64.4% here against 89.0% there, and OSWorld named 2.1 here and 2.0 there — and both are disclosed on Claude Opus 5.5's ## Conflicting Reports. The Chartography gap lines up with this table's no tools label against that one's with tools.
No per-effort absolute figure is published. The page asserts that on several benchmarks Sonnet 5.5 at Low or Medium effort beats Sonnet 5's best score for about a tenth of the cost per task, and renders the comparison as charts rather than a table, so the individual points are not quotable.
Use Cases
Anthropic's stated positioning: well-scoped everyday tasks, fixing bugs, and creating polished documents, slides and spreadsheets — the faster, lower-cost complement to Opus 5.5, which it reserves for complex work requiring careful judgment (source).
Two capability claims sit outside coding:
- Computer use. OSWorld 2.1 80.1% (partial), against Sonnet 5's 57.0% — a 23.1-point gain on the tier's weakest axis, and within 1.7 of Opus 5.5.
- Screenshot-only long-horizon play. "it's the first Sonnet model to beat Pokémon Red working only from screenshots" — the qualifier is the claim, since the feat itself is not new to the Claude line, only to this tier and to this input channel.
The docs' default-model recommendation is unchanged by the release: Opus 5.5 for most workloads, Claude Fable 5.1 for demanding reasoning and long-horizon agentic work. Sonnet 5.5 is not positioned as anyone's default.
Compared To
- Claude Sonnet 5 — direct predecessor, same $2/$10 rate card. Beaten on all eight published rows, and by a wide margin on the three that changed harness version or measure tools (Terminal-Bench 4.0, Chartography, OSWorld 2.1).
- Claude Opus 5.5 — released six days earlier at $4/$20. Default
effort
medium, where Sonnet 5.5's ishigh(docs), so the two models' headline numbers are not produced at the same setting and Eval Harness Configuration holds why that matters. - Claude Fable 5.1 — remains the line's reasoning and long-horizon model; its cache reads are priced at 2.5% of base input against Sonnet 5.5's 10%.
Sources
- Introducing Claude Sonnet 5.5 (snapshot)
- Claude models overview — context window, max output, pricing, model IDs, cutoffs, default effort, retirement date