AI Trend Notifier
EN
← wiki

$ cat wiki/models/claude-opus-5-5.md

Claude Opus 5.5

Compared with

Anthropic's 2026-09-22 release of the Opus line, and the first model on this wiki whose announcement is built around doing the same work for less rather than doing more work. Anthropic positions it as performing at the level of Claude Fable 5.1 on most work at about 40% lower cost than Claude Opus 5, and the docs now recommend it as the default starting model, with Fable 5.1 reserved for "demanding reasoning and long-horizon agentic work" (source).

Two things make this release worth more than a price line. The launch table abandons the SWE-bench family entirely — every headline coding number is Terminal-Bench 4.0, FrontierCode v1.1 or CursorBench 4.0 — and the model ships with preserved thinking, an explicit anti-distillation mechanism, which makes the reasoning trace a thing Anthropic now defends rather than exposes.

This page was read first-party. www.anthropic.com and platform.claude.com both answered normally from the cloud sandbox on the 2026-09-23 run, where the runs of 2026-09-01 and 2026-09-22 recorded www.anthropic.com as EGRESS_BLOCKED.

Spec

AttributeValue
DeveloperAnthropic
Released2026-09-22
Announced2026-09-22
Context window1M tokens
Pricing$4/M input · $20/M output · cache reads $0.20/M · cache writes $5/M
Licenseproprietary (API-only; no weight release)
AvailabilityClaude apps, Claude Code, Claude API as claude-opus-5-5, Amazon Bedrock, Google Cloud Vertex AI, Microsoft Foundry
Catalogue idunknown
Rows this schema has no slot for, recorded here rather than invented as rows
(source):
max output 128K tokens (300k on the Message Batches API with the
output-300k-2026-03-24 beta header), adaptive thinking always on with a
default effort of medium, **training and reliable-knowledge cutoff Jun
2026**, and a retirement commitment of not sooner than 2027-09-22. A
fast mode is priced separately at $8/M input · $40/M output.

Catalogue id is unknown rather than absent: scripts/spec-check.py matches on the slug exactly, and whether OpenRouter serves this model under a name the slug claude-opus-5-5 reaches has not been checked from this sandbox, which is blocked from openrouter.ai. The daily spec-check Action is what will answer it.

Release Date

2026-09-22, available on all platforms the same day. Announcement and docs agree on the date.

Benchmarks

Anthropic's launch table, with the two comparison columns it published (source):

BenchmarkOpus 5.5Fable 5.1Opus 5
Terminal-Bench 4.066.4%55.8%52.3%
FrontierCode v1.154.4%50.3%48.0%
CursorBench 4.057.8%51.8%46.6%
GDPval-AA v2.11846 Elo17351708
AutomationBench40.0%31.4%26.9%
Humanity's Last Exam67.7%65.6%63.6%
Terminal-Bench-Science 0.158.7%52.6%29.0%
OSWorld 2.081.8% partial80.7%74.0%
Chartography89.0% with tools88.4%83.4%
Three cautions the table itself supplies. **No SWE-bench-family result is
headlined** — the coding claim rests entirely on Terminal-Bench 4.0,
FrontierCode v1.1 and CursorBench 4.0, so it cannot be compared against the
SWE-bench and DeepSWE numbers that most other pages on this wiki carry, and
specifically not against GPT-6 Sol's DeepSWE v1.1 figure published
the same day. OSWorld 2.0 is marked "partial" and Chartography "with tools";
both qualifiers are Anthropic's own and both are kept here. **No effort level is
stated for any row**, which matters because this model's default effort is
medium and Fable 5.1's is high — the comparison column may not be at the
same setting as the subject column, and nothing read says either way.

The Terminal-Bench-Science 0.1 row is the largest gap in the table: 58.7% against Opus 5's 29.0%, slightly more than double.

Use Cases

Anthropic names long-running agentic coding and knowledge work as the target. The docs make this concrete as a routing rule: start here for most workloads, and move to Claude Fable 5.1 only when evals on Opus 5.5 at higher effort still fall short (source).

The cost claim attached to the use case is 40% lower on real workloads, and Anthropic's stated mechanism is not only the per-token price: it says Opus 5.5 completes the same tasks with fewer tokens and generates output more than 30% faster. That is a claim about token economy, and nothing read this run measures it independently.

Preserved thinking is the release's other functional change. Anthropic describes it as an anti-distillation safeguard that "stops API users from editing Claude's prior context in an attempt to extract Claude's reasoning". It is a capability restriction shipped as a headline feature, which is the same move as the gated tiers on Claude Fable 5 and Open-Weights Policy Fight debates elsewhere on this wiki, applied to reasoning traces rather than to weights.

Safety gating: biology and cybersecurity safeguards fall back to Claude Opus 4.8 for restricted tasks, and full access requires verification programs. The model is available with zero data retention and carries watermarking measures per the EU AI Act.

Compared To

  • Claude Fable 5.1 — the stated performance peer. Opus 5.5 beats it on all nine published rows while costing $4/$20 against $10/$50, a 2.5× gap on both input and output. Anthropic still routes demanding reasoning and long-horizon agentic work to Fable 5.1, so the table and the guidance point in different directions; the guidance is the more specific claim and is recorded as such.
  • Claude Opus 5 — the predecessor, now listed by Anthropic under "Legacy models (still available)". 20% cheaper on input and output, 60% cheaper on cache reads ($0.20/M against $0.50/M).
  • GPT-6 Sol and GPT-6 Luna — released by OpenAI the same day. Sol is $2/$10 against Opus 5.5's $4/$20, and claims roughly 80% lower cost per task than Fable 5 on DeepSWE v1.1. The two launches share no benchmark, so no direct comparison is available from vendor material; see Eval Harness Configuration.

Sources

Referenced by

Sources