$ cat wiki/models/gpt-5-6-cyber.md
GPT-5.6-Cyber
Spec
| Attribute | Value |
|---|---|
| Developer | OpenAI |
| Released | 2026-08-10 |
| Announced | 2026-08-10 |
| Context window | unknown |
| Pricing | unknown |
| License | proprietary |
| Availability | Daybreak Red tier only — vetted access for authorized vulnerability research, exploit validation and security testing; also via Amazon Bedrock from 2026-08-11 (bedrock-mantle endpoint), same vetting |
| Built on top of Sol, trained to improve at finding | |
| zero-day vulnerabilities and developing exploit chains, and explicitly to | |
| refuse less on higher-risk dual-use cyber work | |
| (source). |
No published price and no system card in anything read — a day after release, for a model whose stated purpose is offensive-capability work. That absence is recorded as the finding rather than as a research gap.
Release Date
2026-08-10, alongside the two-tier restructuring of Daybreak (source).
2026-08-11 — Daybreak on AWS. One day later, OpenAI published "Daybreak models
are now available on AWS": both tiers are reachable through Amazon Bedrock,
via the Bedrock console or the Responses API on the bedrock-mantle
endpoint, inside customers' own AWS security and governance workflows
(source).
The vetting gate is unchanged — the terms carried are "eligible" and "once approved". This is a distribution channel added to an access-controlled programme, not a loosening of it, and it extends an existing arrangement: OpenAI's frontier models and Codex already reached general availability on AWS earlier in 2026. Still no published price, on either channel.
Benchmarks
One figure was published, and it measures compliance rather than capability. OpenAI's Advanced Cybersecurity Completion Rate — how often the model completes rather than refuses this class of task:
| Model / configuration | Completion rate |
|---|---|
| GPT-5.6-Cyber | 95.0% |
| GPT-5.5-Cyber (predecessor) | 57.3% |
| Sol via Daybreak Blue | 2.0% |
| Sol with standard safeguards | 1.5% |
| (source) |
Read the bottom two rows together: Daybreak Blue moves Sol from 1.5% to 2.0%. The "loosened" general-purpose tier is a half-point change, and essentially all of the distance to 95.0% is in the purpose-trained model, not in relaxing a safeguard. That is a meaningful distinction and it is OpenAI's own number.
No capability benchmark was published — nothing on exploit quality, false positive rate, or performance against a human red team. The single claimed result is anecdotal: since training finished, OpenAI reports using the model to investigate V8, the JavaScript engine in Chrome, and uncovering two previously unknown vulnerabilities that could be chained to corrupt memory and escape the V8 heap sandbox (source).
For scale against the one comparable figure this wiki holds, Gemini 3.5 Flash Cyber was reported to have found 55 Chrome V8 vulnerabilities in its restricted pilot. The two are not comparable — no source states the time window, the methodology or the severity in either case — and they are placed together only because V8 is the same target.
Use Cases
- Authorized vulnerability research, exploit validation and security testing, the three uses named by the Daybreak Red gate (source).
- Eric Wallace states OpenAI is using it internally across defensive work (source).
Access is the product as much as the model is. Daybreak now has two tiers:
| Tier | What it carries | Gate |
|---|---|---|
| Daybreak Blue | Frontier general-purpose models, including GPT-5.6 Sol (and Terra, Luna), with loosened cyber safeguards | Approved defenders, everyday security work |
| Daybreak Red | GPT-5.6-Cyber | Tighter vetting — authorized research, validation, testing |
| Access depends on identity verification, account security, monitoring, | ||
| approved-use restrictions and legal attestations, and every individual | ||
| Daybreak account must adopt a hardware security key from 2026-09-01 | ||
| (source). |
Compared To
| Model | Access | What is gated | Published capability figure |
|---|---|---|---|
| GPT-5.6-Cyber | Daybreak Red, vetted | the model itself | none |
| GPT-5.6 Sol (and Terra, Luna) via Daybreak Blue | vetted defenders | the safeguards | completion rate 2.0% |
| Gemini 3.5 Flash Cyber | restricted pilot, governments and trusted partners | the model itself | 55 Chrome V8 vulns |
| Astra | not released — development slowed | everything | none |
| The pattern across the first three is the same answer to the same problem: | |||
| **a cyber-capable model exists, so restrict who holds it rather than what it will | |||
| do.** The fourth row is the one that does not fit, and it is OpenAI's own. |
Conflicting Reports
No source disagrees with another on the figures. The tension recorded below is between two of OpenAI's own decisions, not between two accounts of one.
Open Questions
- Three days separate a pause and a release. On 2026-08-07 OpenAI stated it cannot rule out the Critical cyber threshold for Astra under the Preparedness Framework and slowed development. On 2026-08-10 it shipped a model whose headline figure is an Advanced Cybersecurity Completion Rate of 95.0% against its predecessor's 57.3% on the same class of work. Outlets framed the sequence as a reversal — TheNextWeb: "OpenAI paused a model over cyber risk on Friday. On Monday it shipped one trained to refuse less"; SQ Magazine: "OpenAI Lifts GPT-5.6 Cyber Guardrails Days After Astra Halt". The two decisions are not necessarily inconsistent — one gates an unreleased frontier model's development, the other gates distribution of a narrower model behind vetting — but no source read states that OpenAI addressed the relationship, and the framework's tier vocabulary gives no answer either (source).
- What is GPT-5.6-Cyber's Preparedness assessment? Astra's is the reason it was slowed. Nothing read gives a tier for the model that shipped.
- No context window, price, parameter count or system card.
- What happens if a vetted account is compromised? The hardware-key requirement from 2026-09-01 implies the question was asked; no answer was published.
- The 95.0% figure is a completion rate on OpenAI's own internal set. Is the set published? Nothing read says so.
Sources
- OpenAI — Expanding Daybreak as the Cyber Defense Window Narrows (source)
- OpenAI — Daybreak
- OpenAI — Putting frontier cyber models in more trusted hands
- Unite.AI — OpenAI Expands Daybreak With Two Tiers and a New Cybersecurity Model
- CNBC — OpenAI expands Daybreak cybersecurity initiative as AI agent threats evolve
- TheNextWeb — OpenAI ships GPT-5.6-Cyber, a model trained to refuse less
- eesel AI — GPT-5.6-Cyber: what it is and who can actually get it