AI Trend Notifier
EN한
← wiki

$ cat wiki/models/gemini-4-argon.md

Gemini 4 Argon

Compared with

Google DeepMind's 2026-09-30 frontier release, announced by SVP Koray Kavukcuoglu, and the first named model of the Gemini 4 generation — whose pre-training run this wiki has carried since July on Gemini 4 (source).

Two things make it unusual rather than just large. The first is the output budget: 1M output tokens against the Gemini line's prior 64K, a 15× jump in what the model is allowed to emit, paired with a 2M-token context window. The second is who gets it. Argon ships first to vetted cyber defenders through the Fairwind Program — and to them "without cyber guardrails" — with paid API customers and Google AI Ultra subscribers named as next and no date given. A frontier model whose launch cohort is a security programme rather than an API is a different kind of release from the nine Gemini pages above it.

This page was not read first-party. deepmind.google, blog.google, allenai.org, www.latent.space and www.marktechpost.com all answered EGRESS_BLOCKED at CONNECT from the cloud sandbox on 2026-10-02, so every figure here rests on three independent search passes and records only what two or more stated identically. Where the passes disagreed — notably on how many rows of Google's own comparison table Argon leads — nothing is adopted. See ## Conflicting Reports.

Spec

AttributeValue
DeveloperGoogle DeepMind
Released2026-09-30 (Fairwind Program cohort only)
Announced2026-09-30
Context window2M tokens
Pricing$2/M input · $10/M output · cached input $0.10/M (introductory; stated to double to $4/$20)
Licenseproprietary
AvailabilityFairwind Program (trusted cyber defenders) only; paid API and Google AI Ultra stated as next, no date
Catalogue idunknown
Rows with no slot in this schema: maximum output 1,000,000 tokens, stated as
up from 64,000 in prior Gemini models
(source).

Catalogue id is unknown and is expected to stay that way for now: as of 2026-09-30 there is no published API model ID, and the model is absent from the OpenRouter and models.dev catalogues and from the Vertex AI, Gemini CLI, Cursor and GitHub Copilot model docs. openrouter.ai is blocked from this sandbox in any case; the daily spec-check Action is what will report a match once one exists.

Release Date

  • Announced and released to the Fairwind cohort: 2026-09-30.
  • Paid API customers and Google AI Ultra subscribers: stated as next, no date announced.
  • Google stated it is strengthening four safeguards before a wider release. The four are not enumerated in anything read — which makes the gate on general availability unverifiable from outside.

Benchmarks

Every figure below is Google's own, via search passes. No independent verification of any of them has been published, and the comparison table was not read.

BenchmarkGemini 4 ArgonBest competitor named
DeepSWE v1.177.9% (stated SOTA)Claude Opus 5.5 74.2% · Astra 74.1%
FrontierSWE v255.0%Astra 65.5% (Argon −10.5)
Terminal-bench 4.057.4% (last of four)Claude Opus 5.5, ~9 points ahead
CWE-bench v1tieAstra (tie)
PostTrainBenchlossClaude Opus 5.5
OSWorld-2.0loss — no figure statedunknown
The shape of this table is worth more than the DeepSWE headline. **Argon takes
the new long-horizon SWE benchmark and loses the two agentic-terminal ones** —
FrontierSWE v2 to Astra by 10.5 points, Terminal-bench 4.0 to Opus 5.5 by about
nine. A model announced for "long-horizon software engineering" is **last of four
on the benchmark that measures a terminal agent**, which is a result Google
published rather than one anybody extracted from it.

The OSWorld-2.0 figure is missing from all three passes and is recorded as missing rather than estimated.

Use Cases

Google names three, in this order (source):

  • Long-horizon software engineering.
  • Enterprise knowledge work — specifically legal drafting and financial research, both described as long, multi-step jobs.
  • Cybersecurity defence. The model is stated to autonomously find, validate and patch critical software vulnerabilities.

The cyber framing is not a side note; it determines the release. The Fairwind Program is stated to have had more than 650 participating partners globally at its own launch — national cybersecurity authorities, critical infrastructure operators, technology companies and security vendors — and it is that cohort, plus Google's internal teams, that receives the model "without cyber guardrails". Misuse defences for cyber and chemical, biological, radiological and nuclear attacks are stated to have been tested by internal and external red teams.

This lands one day after Anthropic published measured ExploitBench figures for an open-weight competitor, and the two together are recorded on AI-Enabled Cyberattacks: one lab publishing a rival's cyber capability as a number, the other shipping its own cyber capability to defenders with the guardrails removed.

Compared To

  • GPT-6.1 Sol — the price comparison is exact and almost certainly deliberate. Sol shipped on 2026-09-29 at $2/M input · $10/M output · cached $0.10/M; Argon's introductory rate the next day is the same three numbers. Argon offers 2M context against Sol's 1.05M and 1M output against Sol's 128K. Argon's rate is stated to double to $4/$20, Sol's is not described as introductory — so the matching is a launch posture, not a standing price.
  • Astra — still ahead on FrontierSWE v2 (65.5% vs 55.0%) and tied on CWE-bench v1.
  • Claude Opus 5.5 — still ahead on Terminal-bench 4.0 and PostTrainBench; behind on DeepSWE v1.1 (74.2% vs 77.9%).
  • Gemini 4 — the July 2026 pre-training announcement this model realises. That page still reads Released: not yet; see its own note.
  • Gemini 3.8 Flash and the rest of the Gemini 3.x line — the generation Argon supersedes at the top.

Sources

Conflicting Reports

  • How many benchmark rows Google claims Argon leads — unresolved, and no count is adopted. The three search passes give "13 of 19 rows led outright", "tops Astra and Opus 5.5 in 12 of 18 tests", and "beats GPT-6 Astra on 14 of 19 rows and ties it on CWE-bench v1" alongside "leads 13 of 18 launch benchmarks". None of them quotes the table. Two different denominators (18 and 19) are in play, which suggests at least one pass is counting a different table or dropping the tied row. This wiki states the per-benchmark figures it can attribute and no aggregate. Resolving it requires reading Google's page, which this sandbox cannot reach (source).

Referenced by

Sources