$ cat wiki/models/gemini-4-argon.md
Gemini 4 Argon
Compared with
- GPT-6.1 Sol
- Claude Sonnet 5.5
- MiMo-V2.6-Pro
- Grok 4.7
- Ternary Bonsai 2 27B
- Fugu Max
- Kimi K2.8 Preview
- DeepSeek V4.1-Flash
- K2 Horizon
- Muse Spark 1.3
- Hy4 preview
- GLM-5.3-Flash
- Granite 4.2
- Ling-3.0-tiny
- Laguna S 2.1
- Inkling
- LongCat-2.0
- MiniMax M3
- Claude Opus 5.5
- GPT-6 Luna
- GPT-6 Sol
- Gemini 3.8 Live
- Fugu Ultra v2
- GPT-Image-2.5 Flare
- GPT-Image-2.5 Sunburst
- Astra
- Gemini 3.8 Flash
- Claude Fable 5.1
- GLM-5.3
- Qwen 3.8 27B
- DeepSeek V4-Pro-0813
- Gemini 3.7 Flash
- Muse Glimmer
- Grok 4.6
- Muse Spark 1.2
- Qwen 3.8 Max
- DeepSeek V4-Flash
- Claude Opus 5
- Gemini 3.5 Flash-Lite
- Gemini 3.6 Flash
- DeepSeek V4
- Kimi K3
- GPT-5.6 Sol
- Grok 4.5
- Claude Sonnet 5
- GLM-5.2
- Claude Fable 5
- Claude Opus 4.8
- Grok Build
- Claude Opus 4.7
- Muse Spark
Google DeepMind's 2026-09-30 frontier release, announced by SVP Koray Kavukcuoglu, and the first named model of the Gemini 4 generation — whose pre-training run this wiki has carried since July on Gemini 4 (source).
Two things make it unusual rather than just large. The first is the output budget: 1M output tokens against the Gemini line's prior 64K, a 15× jump in what the model is allowed to emit, paired with a 2M-token context window. The second is who gets it. Argon ships first to vetted cyber defenders through the Fairwind Program — and to them "without cyber guardrails" — with paid API customers and Google AI Ultra subscribers named as next and no date given. A frontier model whose launch cohort is a security programme rather than an API is a different kind of release from the nine Gemini pages above it.
This page was not read first-party. deepmind.google, blog.google,
allenai.org, www.latent.space and www.marktechpost.com all answered
EGRESS_BLOCKED at CONNECT from the cloud sandbox on 2026-10-02, so every figure
here rests on three independent search passes and records only what two or
more stated identically. Where the passes disagreed — notably on how many rows
of Google's own comparison table Argon leads — nothing is adopted. See
## Conflicting Reports.
Spec
| Attribute | Value |
|---|---|
| Developer | Google DeepMind |
| Released | 2026-09-30 (Fairwind Program cohort only) |
| Announced | 2026-09-30 |
| Context window | 2M tokens |
| Pricing | $2/M input · $10/M output · cached input $0.10/M (introductory; stated to double to $4/$20) |
| License | proprietary |
| Availability | Fairwind Program (trusted cyber defenders) only; paid API and Google AI Ultra stated as next, no date |
| Catalogue id | unknown |
| Rows with no slot in this schema: maximum output 1,000,000 tokens, stated as | |
| up from 64,000 in prior Gemini models | |
| (source). |
Catalogue id is unknown and is expected to stay that way for now: as of
2026-09-30 there is no published API model ID, and the model is absent from
the OpenRouter and models.dev catalogues and from the Vertex AI, Gemini CLI,
Cursor and GitHub Copilot model docs. openrouter.ai is blocked from this
sandbox in any case; the daily spec-check Action is what will report a match
once one exists.
Release Date
- Announced and released to the Fairwind cohort: 2026-09-30.
- Paid API customers and Google AI Ultra subscribers: stated as next, no date announced.
- Google stated it is strengthening four safeguards before a wider release. The four are not enumerated in anything read — which makes the gate on general availability unverifiable from outside.
Benchmarks
Every figure below is Google's own, via search passes. No independent verification of any of them has been published, and the comparison table was not read.
| Benchmark | Gemini 4 Argon | Best competitor named |
|---|---|---|
| DeepSWE v1.1 | 77.9% (stated SOTA) | Claude Opus 5.5 74.2% · Astra 74.1% |
| FrontierSWE v2 | 55.0% | Astra 65.5% (Argon −10.5) |
| Terminal-bench 4.0 | 57.4% (last of four) | Claude Opus 5.5, ~9 points ahead |
| CWE-bench v1 | tie | Astra (tie) |
| PostTrainBench | loss | Claude Opus 5.5 |
| OSWorld-2.0 | loss — no figure stated | unknown |
| The shape of this table is worth more than the DeepSWE headline. **Argon takes | ||
| the new long-horizon SWE benchmark and loses the two agentic-terminal ones** — | ||
| FrontierSWE v2 to Astra by 10.5 points, Terminal-bench 4.0 to Opus 5.5 by about | ||
| nine. A model announced for "long-horizon software engineering" is **last of four | ||
| on the benchmark that measures a terminal agent**, which is a result Google | ||
| published rather than one anybody extracted from it. |
The OSWorld-2.0 figure is missing from all three passes and is recorded as missing rather than estimated.
Use Cases
Google names three, in this order (source):
- Long-horizon software engineering.
- Enterprise knowledge work — specifically legal drafting and financial research, both described as long, multi-step jobs.
- Cybersecurity defence. The model is stated to autonomously find, validate and patch critical software vulnerabilities.
The cyber framing is not a side note; it determines the release. The Fairwind Program is stated to have had more than 650 participating partners globally at its own launch — national cybersecurity authorities, critical infrastructure operators, technology companies and security vendors — and it is that cohort, plus Google's internal teams, that receives the model "without cyber guardrails". Misuse defences for cyber and chemical, biological, radiological and nuclear attacks are stated to have been tested by internal and external red teams.
This lands one day after Anthropic published measured ExploitBench figures for an open-weight competitor, and the two together are recorded on AI-Enabled Cyberattacks: one lab publishing a rival's cyber capability as a number, the other shipping its own cyber capability to defenders with the guardrails removed.
Compared To
- GPT-6.1 Sol — the price comparison is exact and almost certainly deliberate. Sol shipped on 2026-09-29 at $2/M input · $10/M output · cached $0.10/M; Argon's introductory rate the next day is the same three numbers. Argon offers 2M context against Sol's 1.05M and 1M output against Sol's 128K. Argon's rate is stated to double to $4/$20, Sol's is not described as introductory — so the matching is a launch posture, not a standing price.
- Astra — still ahead on FrontierSWE v2 (65.5% vs 55.0%) and tied on CWE-bench v1.
- Claude Opus 5.5 — still ahead on Terminal-bench 4.0 and PostTrainBench; behind on DeepSWE v1.1 (74.2% vs 77.9%).
- Gemini 4 — the July 2026 pre-training announcement this model
realises. That page still reads
Released: not yet; see its own note. - Gemini 3.8 Flash and the rest of the Gemini 3.x line — the generation Argon supersedes at the top.
Sources
- Captured snapshot — including the three search queries used and the disputed figures
- blog.google — Gemini 4 Argon (blocked, not read)
- deepmind.google mirror (blocked, not read; prefetch candidate #41)
Conflicting Reports
- How many benchmark rows Google claims Argon leads — unresolved, and no count is adopted. The three search passes give "13 of 19 rows led outright", "tops Astra and Opus 5.5 in 12 of 18 tests", and "beats GPT-6 Astra on 14 of 19 rows and ties it on CWE-bench v1" alongside "leads 13 of 18 launch benchmarks". None of them quotes the table. Two different denominators (18 and 19) are in play, which suggests at least one pass is counting a different table or dropping the tied row. This wiki states the per-benchmark figures it can attribute and no aggregate. Resolving it requires reading Google's page, which this sandbox cannot reach (source).