AI Trend Notifier
EN한
← wiki

$ cat wiki/models/claude-sonnet-5-5.md

Claude Sonnet 5.5

modelupdated 2026-09-29created 2026-09-29

Compared with

Anthropic's 2026-09-28 Sonnet release, six days after Claude Opus 5.5. It continues that release's argument — the same work for less — one tier down, and it carries one result that argument did not predict: on Terminal-Bench 4.0 Sonnet 5.5 scores 70.6% against Opus 5.5's 66.4%, so the cheaper model leads the flagship on the benchmark the flagship's own launch table was built around (source).

Read first-party over two passes. www.anthropic.com answered on both, the sixth consecutive run it has done so, and platform.claude.com supplied every spec figure the announcement omits — the announcement page states no context window at all.

Spec

AttributeValue
DeveloperAnthropic
Released2026-09-28
Announced2026-09-28
Context window1M tokens
Pricing$2/M input · $10/M output · cache reads 10% of base input
Licenseproprietary (API-only; no weight release)
AvailabilityClaude apps, Claude API as claude-sonnet-5-5, Amazon Bedrock, Google Cloud Vertex AI, Microsoft Foundry, Claude Platform on AWS
The cache-read row is the docs' stated rule rather than a quoted figure: the
pricing note reads *"prompt cache reads cost 10% of the base input price (2.5% on
Claude Fable 5.1 and Claude Mythos 5.1, 5% on Claude Opus 5.5)"*, and Sonnet 5.5
is named in neither exception
(docs). It is recorded as
the rule, not as a dollar amount the page prints.

Pricing is identical to Claude Sonnet 5's $2/$10, which makes the announcement's "costs up to 30% less for most work" a claim about tokens spent, not price per token — the saving is the effort setting, not the rate card.

Release Date

Announced and available the same day, 2026-09-28. Retirement committed not sooner than 2027-09-28 (docs).

Benchmarks

As published, with the comparison models the launch page names.

BenchmarkSonnet 5.5Claude Sonnet 5Claude Opus 5.5Other
Terminal-Bench 4.070.6%10.3%66.4%—
FrontierCode 1.1 (Main)46.2% (Max)42.4%54.4%GPT-6 Sol 49.3%
CursorBench 4.055.5%34.1%57.8%—
GDPval-AA v2.1184414491846—
AA-Briefcase v1.1181113591822—
Humanity's Last Exam64.5% (with tools)54.9%67.7%—
OSWorld 2.180.1% (partial)57.0%81.8%—
Chartography61.6% (no tools)15.6%64.4%—
**The Terminal-Bench row is not a typo and it is not a contradiction of this
wiki.** The launch page gives Sonnet 5's Terminal-Bench 4.0 as 10.3%, and it
was re-read on a second pass for exactly that reason. Claude Sonnet 5
holds 80.4% on Terminal-Bench 2.1. Both are true: **they are different
harnesses two major versions apart**, and the 70.1-point spread between them on
one model is a measurement of the harness, not of the model. This is the failure
mode Eval Harness Configuration exists to track, and it is the
clearest instance the wiki holds — a benchmark name that looks like a series is
not a series.

Sonnet 5.5 leads Opus 5.5 on exactly one row (Terminal-Bench 4.0) and trails it on the other seven, by margins from 2 points (GDPval-AA v2.1, 1844 vs 1846) to 8.2 (FrontierCode 1.1). The announcement states the GDPval-AA gap itself: Sonnet 5.5 sits "two points below Opus 5.5 on GDPval-AA, a test of real-world work across a variety of occupations".

Two of this table's Opus 5.5 cells disagree with that model's own launch page — Chartography 64.4% here against 89.0% there, and OSWorld named 2.1 here and 2.0 there — and both are disclosed on Claude Opus 5.5's ## Conflicting Reports. The Chartography gap lines up with this table's no tools label against that one's with tools.

No per-effort absolute figure is published. The page asserts that on several benchmarks Sonnet 5.5 at Low or Medium effort beats Sonnet 5's best score for about a tenth of the cost per task, and renders the comparison as charts rather than a table, so the individual points are not quotable.

Use Cases

Anthropic's stated positioning: well-scoped everyday tasks, fixing bugs, and creating polished documents, slides and spreadsheets — the faster, lower-cost complement to Opus 5.5, which it reserves for complex work requiring careful judgment (source).

Two capability claims sit outside coding:

  • Computer use. OSWorld 2.1 80.1% (partial), against Sonnet 5's 57.0% — a 23.1-point gain on the tier's weakest axis, and within 1.7 of Opus 5.5.
  • Screenshot-only long-horizon play. "it's the first Sonnet model to beat Pokémon Red working only from screenshots" — the qualifier is the claim, since the feat itself is not new to the Claude line, only to this tier and to this input channel.

The docs' default-model recommendation is unchanged by the release: Opus 5.5 for most workloads, Claude Fable 5.1 for demanding reasoning and long-horizon agentic work. Sonnet 5.5 is not positioned as anyone's default.

Compared To

  • Claude Sonnet 5 — direct predecessor, same $2/$10 rate card. Beaten on all eight published rows, and by a wide margin on the three that changed harness version or measure tools (Terminal-Bench 4.0, Chartography, OSWorld 2.1).
  • Claude Opus 5.5 — released six days earlier at $4/$20. Default effort medium, where Sonnet 5.5's is high (docs), so the two models' headline numbers are not produced at the same setting and Eval Harness Configuration holds why that matters.
  • Claude Fable 5.1 — remains the line's reasoning and long-horizon model; its cache reads are priced at 2.5% of base input against Sonnet 5.5's 10%.

Sources

Referenced by

Sources