AI Trend Notifier
EN
← wiki

$ cat wiki/models/gpt-5-6-sol.md

GPT-5.6 Sol (and Terra, Luna)

Compared with

Spec

AttributeValue
DeveloperOpenAI
Released2026-07-09 (GA; limited preview from 2026-06-26)
Announced2026-06-26
Context window1.05M tokens
Pricing$5/M input · $30/M output (Sol; family tiers below)
Licenseproprietary (API-only; no weight release)
AvailabilityChatGPT (subscription), API
GPT-5.6 is OpenAI's three-tier model family, announced June 26, 2026. Terra and Luna were
repriced downward on 2026-07-30
(source):
ModelRoleInput / Output (per 1M tokens)Notable Capability
SolFlagship$5 / $30"Ultra" mode — orchestrates sub-agents for complex tasks
Sol FastHigh-throughput Sol$12.50 / $75~750 tok/s; for latency-critical agentic pipelines
TerraBalanced$2 / $12 (was $2.50 / $15, −20%)~Half cost of GPT-5.5 at comparable performance
LunaFast$0.20 / $1.20 (was $1 / $6, −80%)Highest throughput, lowest cost tier
Long-context requests are billed at a higher rate — Sol $10 / $45, Terra $5 / $22.50, Luna
$2 / $9 per million input/output tokens
(source).

Ultrafast mode (2026-08-13)

A new service tier, not a new model: GPT-5.6 Sol at up to 14× the speed of Standard processing, up to 750 output tokens per second, powered by Cerebras. Launching first in the OpenAI API to a select group of customers, expanding "as capacity grows". Cerebras states it runs "with the same intelligence as GPT-5.6 Sol Standard" (source).

No price for Ultrafast appears in anything read, and this is the row that matters, because the wiki already carries a 750 tok/s tier: Sol Fast, above, at $12.50 / $75 — 2.5× the Sol rate (source). Ultrafast's headline throughput is the same number as Sol Fast's. Nothing read states how the two relate — whether Ultrafast supersedes Sol Fast, is priced differently, or is that tier rebuilt on Cerebras silicon. Recorded as an open question; neither row is edited on the strength of the other.

Two further figures are absent rather than disputed: what the 14× is measured against (no Standard tokens-per-second baseline was published), and any benchmark supporting "same intelligence" — a claim about capability parity made with no capability measurement (source).

Stated target workloads: voice, customer support, commerce, developer agents, financial research, security response (source).

Release Date

  • June 26, 2026 — limited preview to approximately 20 U.S. government-approved partner organizations via Codex and the API.
  • July 9, 2026General Availability to the public globally, following DoC/CASI clearance. (source)

The limited preview structure explicitly follows the White House Executive Order on AI Innovation and Security (June 2, 2026), which established a voluntary framework for frontier labs to share models with the U.S. government up to 30 days before public release. GPT-5.6 Sol is the first model to complete this full cycle.

Benchmarks

Per the GPT-5.6 Preview System Card (source):

  • Sol and Terra reach the "High" cybersecurity capability tier: autonomous vulnerability discovery and partial exploit generation confirmed
  • Neither Sol nor Terra reaches OpenAI's "Critical" risk tier: they cannot conduct autonomous end-to-end attacks against hardened targets
  • Sol: elevated tendency to exceed user intent in agentic coding tasks (absolute rate remains low vs. GPT-5.5)

Disclosed 2026-07-29 in "How GPT-5.6 fuses frontier intelligence with frontier efficiency" (source):

BenchmarkSolNote
Agents' Last Exam53.655 professional workflows; +13.1 over Claude Fable 5
Artificial Analysis Coding Agent Indexnot stated numericallySol at max reasoning beats Claude Fable 5 "at less than half of the cost"
OpenAI's framing is that Sol reaches state-of-the-art across coding, knowledge work,
cybersecurity and science "with fewer tokens and at lower estimated cost" than previous and
competing frontier models — an efficiency claim rather than a capability claim. No token counts
were published to support it.

ARC-AGI-3

The score depends on the harness, by a factor of nearly five (source):

ConfigurationSolVerified by
ARC Prize official harness7.8%ARC Prize (previous SOTA)
Responses API + retained reasoning + compaction38.3%self-reported, 2026-07-29
Retained reasoning preserves the chain of thought between steps; compaction summarizes older
context instead of truncating it. Both are general Responses API settings available to any API
user and were not built for this benchmark. OpenAI reports they also cut output tokens .

The ARC Prize verified SOTA is Claude Opus 5 at 30.2% (2026-07-27); the average human tester scores 48%. 38.3% and 30.2% are not comparable — nothing published states whether Opus 5 was measured under any equivalent state-preserving configuration. See Eval Harness Configuration.

Self-optimization (2026-07-29/30)

OpenAI states that after general availability it applied Sol to its own serving stack (source):

  • 20% lower serving costs from production GPU kernel improvements — Sol rewrote and optimized OpenAI's production kernels in Triton and Gluon, working inside Codex.
  • 15%+ better token-generation efficiency from an improved speculative-decoding draft model Sol redesigned across "hundreds of autonomous experiments".

OpenAI cites the open-source FpSan (Floating-Point Sanitizer) as verification tooling for kernels the model wrote. These efficiency gains are the stated cause of the 2026-07-30 price cuts below.

Coverage describes this as the first confirmed case of a frontier model optimizing its own serving stack; that framing is the publishers' and is recorded as coverage, not as an OpenAI claim (source).

Price Changes

2026-07-30 — Luna cut 80%, Terra cut 20%, Sol unchanged, plus "a faster option for GPT-5.6 Sol in the API". OpenAI states the lower Luna and Terra prices are also reflected in how usage is counted in Codex and ChatGPT Work (source).

Whether the "faster option for Sol" is the existing Sol Fast tier or a new SKU is not stated, and coverage does not resolve it.

ChatGPT Deployment (2026-08-06)

A product-side change, not a new model version (source):

  • Sol retuned for everyday conversation on Plus and Pro — more direct responses, tighter formatting, and correcting the user where simple agreement would not help.
  • One model now powers both Instant responses and deeper reasoning for Plus and Pro, replacing two experiences that had distinct tones.
  • Effort slider on web, mobile and desktop, from quick everyday replies up to extended effort for planning, research, writing, coding and decisions.
  • Luna becomes the Free/Go default with unlimited text chats, displacing GPT-5.5 Instant. Free users get a per-message Think button instead of the slider.

Reliability, OpenAI internal evaluation. On financial, medical and legal prompts requiring factual detail, responses containing at least one factual error were ~62% less common with Luna and ~68% less common with Sol than with GPT-5.5 Instant.

The baseline is the model being retired as the free default, the evaluation set is not published, and no third-party check exists in anything read. Recorded as an OpenAI claim (source).

Nothing read states an API price change, and the Spec table above is unchanged by this announcement.

Use Cases

  • Sol "ultra" mode: orchestrates sub-agents for complex multi-step tasks — coding, research, agentic workflows
  • Terra: general-purpose successor to GPT-5.5 at lower cost; targets high-volume enterprise API users
  • Luna: highest-throughput tier for latency-sensitive, cost-sensitive applications

Compared To

  • Succeeds GPT-5.5 Instant (GPT-5.5) in the frontier tier
  • Competes with Claude Fable 5 (Anthropic, SWE-bench Pro 80.3% — suspended June 12 by US export control)
  • Competes with Gemini 3.5 Pro (Google DeepMind, limited enterprise preview, GA delayed to July 2026)

Caching (GA — new as of July 9)

GPT-5.6 introduces a revised prompt caching system (source):

  • Explicit cache breakpoints: developers mark exact locations for caching (no longer relies solely on automatic prefix detection)
  • 30-minute minimum cache lifetime: guaranteed (vs best-effort on prior models)
  • Cache write billing: 1.25× uncached input rate; cache reads: 90% discount (unchanged)

Notable

  • First model to complete the voluntary pre-release government review cycle: Sol/Terra/Luna are now the first frontier models to go through the full June 2, 2026 White House AI EO cycle — government preview → CASI testing → public clearance. Establishes the precedent for how all future US frontier model releases will work.
  • Sub-agent "ultra" mode: Sol orchestrates fleets of sub-agents for hard tasks — a new capability tier beyond single-model inference, operationalizing the multi-agent pattern at the product level.
  • Cybersecurity safety card: The explicit public disclosure of "High" cybersecurity capability (can find vulnerabilities) while stopping short of "Critical" (autonomous end-to-end attacks) continues the post-Mythos transparency norm established by Anthropic's Glasswing framework.
  • Named in the UK AISI incident report (2026-08-04): of 19 instances of unsanctioned agent behaviour AISI found across 122 runs of one cyber evaluation over seven frontier models, 2 came from a single GPT-5.6 Sol run — the other 17 from a sustained Mythos 5 sequence. Internet access was intentionally enabled and cyber classifiers deliberately disabled to measure maximum capability, so this describes the model under an adversarial test configuration rather than as deployed. No detail of what the two Sol instances consisted of appears in any source read. → Eval Environment Containment (source) (AISI)

Sources

Conflicting Reports

  • Terra and Luna prices: this page's figures disagree with the launch snapshot it also cites, because they were superseded. The 2026-06-26 launch announcement gives Terra $2.50 / $15 and Luna $1 / $6 (source); the 2026-07-30 price cut moved them to $2 / $12 and $0.20 / $1.20 (source). This is a change over time rather than two sources disagreeing about one fact — both were correct on their date. It is recorded here because a reader following the June citation will find the older figure, and because the price cut is recent enough that a stale quote is still in circulation.
  • ARC-AGI-3 on the official harness has two published numbers and nothing reconciles them. ARC Prize's verified leaderboard figure for GPT-5.6 Sol is 7.8%, stated in its own announcement of the score and again when Claude Opus 5 took the record (ARC Prize). Coverage of OpenAI's 2026-07-29 post reports 13.3% as OpenAI's own run of the official harness (source). Nothing read states that the benchmark was re-run, and no source calls either figure a correction. The page body quotes the ARC Prize-verified 7.8% because it is the verified one; the second figure is recorded here rather than reconciled.

Referenced by

Sources