AI Trend Notifier
EN
← wiki

$ cat wiki/models/astra.md

Astra

OpenAI's next major model, named publicly for the first time on 2026-08-01 in a post attributing ten solved open problems in mathematics and theoretical computer science to an internal version of it (source).

Six days later it became the first model OpenAI has treated as "Critical" for cybersecurity under its Preparedness Framework, and OpenAI said it is slowing development until safeguards are in place (source).

Spec

AttributeValue
DeveloperOpenAI
Releasednot yet
Announced2026-08-01 (named, no launch)
Context windowunknown
Pricingunknown
Licenseunknown
Availabilityunknown — no API identifier, no availability date stated in any source read
Every unknown above is genuinely unstated: the 2026-08-01 post is a results
announcement, not a product announcement, and no source read here gives a shipping
date, a price, or an endpoint
(source).
The 2026-08-07 post does not fill any of them either — it removes a date rather
than adding one
(source).

Safety Classification

On 2026-08-07 OpenAI published that preliminary internal evaluations of Astra show strong enough agentic coding and cybersecurity performance that it "cannot rule out" the Critical cyber capability level in its own Preparedness Framework, and that it is treating Astra as its first "Critical" model for cybersecurity (source).

Two qualifications belong with that sentence, and OpenAI supplies both: testing is ongoing, and OpenAI states it has not confirmed that Astra crossed the threshold. The published position is that the threshold cannot be ruled out, not that it has been met. Every OpenAI model evaluated before this one — including GPT-5.6 Sol (and Terra, Luna) — was assessed at High rather than Critical.

OpenAI's Critical cyber threshold, as quoted in the post, is a model that can either identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or devise and execute end-to-end novel attack strategies against hardened targets given only a high-level goal.

The announced response:

MeasureAs published
DevelopmentSlowed until safeguards are in place, as the framework requires
Test environmentsIsolated; restricted network and tool access; sandboxed execution
WeightsAdditional protection and encryption of model weights
MonitoringEvery agentic run monitored for risky behaviour; additional controls for agent applications
Safeguard testingRobustness testing scaled up to match the capability level
External testingWith government agencies and selected AI safety organisations before broader deployment
Testing partnersOpenAI to provide recommended security controls to third parties running higher-risk evaluations
Axios reports, as an exclusive, that OpenAI **voluntarily informed the
administration** of its plan to delay
(Axios).
That is reporting, not an OpenAI statement, and is held as such.

What this page cannot say is how much slower. No source read here gives a revised date, a duration, or a condition whose satisfaction would end the delay — which is why Released remains not yet rather than moving to a dated projection.

2026-08-18 — the duration was published, and the pause has ended

Pacing model development in an era of cyber-critical capabilities supplies the number the 08-07 post withheld: the pause lasted a little more than two weeks, and it is over — OpenAI states it assessed the risks, put guardrails in place, and resumed the affected activities (source).

The security controls are restated as ones OpenAI had not previously needed to apply — isolated testing environments, restricted network and tool access, enhanced protection and encryption of model weights, additional monitoring and detection, sandboxed execution. That list matches the 08-07 table above rather than extending it.

What did not change: the Critical designation is neither lifted, confirmed nor revised in anything read, and no release date, price, endpoint or context window is given. Every unknown in the Spec table above stands. So the delay that was "explicitly indefinite" on 08-07 turns out to have been bounded at roughly two weeks for the internal activities it covered — while the thing a reader would take "slowed development" to mean, a shipping date, was never attached to a schedule and still is not.

Scope is disputed in the coverage and is recorded in ## Conflicting Reports below: Fortune's headline says OpenAI "paused AI training for two weeks", while Cryptobriefing reports Sam Altman saying Astra's core training never stopped and that what paused was certain internal activities. The OpenAI extracts themselves say "certain internal activities", which is the reading this page follows.

Release Date

Not released. OpenAI describes the model used as an internal version of Astra, and Sébastien Bubeck (OpenAI) calls it "our next major model" (@SebastienBubeck).

As of 2026-08-07 the schedule is explicitly indefinite: OpenAI says it will slow development "until it has the right safeguards in place" (source).

The cost figure is quoted at Sol API prices rather than at Astra's own — roughly $2,000 in tokens for all ten results — which is consistent with Astra having no published price of its own (source).

Benchmarks

No standard benchmark scores have been published for Astra. What exists instead is a set of ten claimed research results, each shipped with a Lean 4 certificate and a chain-of-thought walkthrough, alongside a 249-page manuscript (source):

ResultField
First explicit non-sofic groupGroup theory
Connes' Rigidity Conjecture disprovedOperator algebras / von Neumann algebras
Quantum parallel repetition for general two-player entangled gamesQuantum complexity
Ehrhart's volume conjecture provedDiscrete geometry
First improvement to the general high-dimensional sphere-packing upper bound since 1978High-dimensional geometry
New circuit complexity lower boundsArithmetic circuit complexity
Monochromatic triangles in multicoloured graphsExtremal combinatorics
Three Erdős problemsCombinatorics
The problems are stated to have been open at least ten years, and in several cases
much longer. Remaining results fall in coding theory and lattice cryptography
(source).

Per Eval Harness Configuration, these are not comparable to benchmark scores and are not recorded as such: there is no shared harness, no baseline, and no other model has been run on the same ten problems.

Use Cases

Only one use is demonstrated: autonomous production of novel mathematical arguments, with human editors organising the output into papers and a separate formalization step producing the Lean certificates (source). See AI for Mathematics for what that division of labour does and does not establish.

Compared To

  • GPT-5.6 Sol (and Terra, Luna) — the currently shipping OpenAI frontier tier, and the price basis the $2,000 figure is quoted in. Sol had already been credited by OpenAI with rewriting its own serving kernels (2026-07-30); Astra is the same programme pointed at research output rather than infrastructure
  • AlphaEvolveGoogle DeepMind's search-based system for mathematical and algorithmic discovery. Different mechanism (evolutionary search over programs against an evaluator) for an overlapping goal
  • Leanstral 1.5Mistral AI's open-weight Lean 4 theorem prover. Occupies the verification half of the pipeline Astra's results depend on, as an open-weight artefact rather than a closed internal one
  • Gemini 3.1 Deep Think — the other frontier line whose public case has leaned on competition-level mathematics
  • Claude Science, Co-Scientist (Google DeepMind) — the same "model as research instrument" framing in the natural sciences

Conflicting Reports

Release status. OpenAI's own framing is an internal version of an unreleased model. Several aggregators instead describe Astra as having been "launched" or "released" — for example KuCoin's flash ("OpenAI has launched a new model, Astra"). No source read here supplies an availability date, price or API identifier, so the page follows OpenAI's framing per the source-priority rule and records the discrepancy rather than resolving it.

What the two-week pause covered. Reporting of the 2026-08-18 post splits on scope, and this page does not resolve it beyond following OpenAI's own wording (source):

ClaimReported by
OpenAI "paused AI training for two weeks"Fortune, 2026-08-18 (headline)
Astra's core training never stopped; what paused was certain internal activities, and new models are still on track to ship soon — attributed to Sam AltmanCryptobriefing
The two are reconcilable if the pause covered a subset of internal activity rather
than the training run, which is what the OpenAI extracts say. **No document read
states the two-week figure and the "core training never stopped" clarification
together**, so neither is adopted over the other.

Relationship to the GPT line. Some coverage glosses Astra as "GPT-6" (@kimmonismus, which writes "Astra model (GPT6?)" with the question mark). Nothing from OpenAI read here states that Astra is or is not a GPT-line successor.

Sources

Referenced by

Sources