$ cat wiki/models/gemini-3-8-flash.md
Gemini 3.8 Flash
Compared with
- Claude Fable 5.1
- Hy4 preview
- GLM-5.3-Flash
- Granite 4.2
- DeepSeek V4-Pro-0813
- Muse Glimmer
- Grok 4.6
- Laguna S 2.1
- Kimi K3
- Inkling
- GPT-5.6 Sol
- LongCat-2.0
- MiniMax M3
- GLM-5.3
- Qwen 3.8 27B
- Gemini 3.7 Flash
- Muse Spark 1.2
- Qwen 3.8 Max
- DeepSeek V4-Flash
- Claude Opus 5
- Gemini 3.5 Flash-Lite
- Gemini 3.6 Flash
- DeepSeek V4
- Grok 4.5
- Claude Sonnet 5
- GLM-5.2
- Claude Fable 5
- Claude Opus 4.8
- Grok Build
- Claude Opus 4.7
- Muse Spark
Spec
| Attribute | Value |
|---|---|
| Developer | Google DeepMind |
| Released | 2026-09-02 |
| Announced | 2026-09-02 |
| Context window | 1,048,576 tokens (65,536 max output) |
| Pricing | $0.75/M input · $3.75/M output (introductory, through 2026-12-31) · $1.50/M · $7.50/M from 2027-01-01 |
| License | proprietary |
| Availability | Gemini API, Google AI Studio, Antigravity, Android Studio, Gemini Enterprise / Gemini Enterprise Agent Platform |
Generally available at announcement, API model id gemini-3.8-flash | |
| (source). |
Every row above except Released and Announced is identical to
Gemini 3.7 Flash — same context window, same output ceiling, same
introductory price, same expiry date, same standard price after it. Three weeks
apart, the only thing the spec table records as having changed is the date
(source).
No parameter count, architecture or weight release is reported, and none is expected on this line.
Release Date
2026-09-02, 20 days after Gemini 3.7 Flash (2026-08-13), which was itself 23 days after Gemini 3.6 Flash (2026-07-21). Three Flash releases in 43 days, each at the same price.
Benchmarks
Vendor-stated, against Gemini 3.7 Flash (source):
| Benchmark | Gemini 3.8 Flash | Gemini 3.7 Flash |
|---|---|---|
| Terminal-Bench 2.1 | 90.8% | 81.6% |
| DeepSWE v1.1 | 73.7% | unknown |
| SWE-bench Pro | 61.6% | 60.4% |
| The three rows are not equally well attested and the page says which is which. | ||
deepmind.google and blog.google are both unreachable from this run's sandbox, | ||
| so the release was captured through three WebSearch passes. Terminal-Bench 2.1 and | ||
| DeepSWE v1.1 were carried by two passes each with identical digits. **SWE-bench Pro | ||
| was carried with digits by one pass only**; the other two corroborated the | ||
| direction — "barely moved", "just over a point" — and no pass contradicted it | ||
| (source). |
The interesting shape is the spread between those rows. Terminal-Bench 2.1 moves 9.2 points and SWE-bench Pro moves 1.2 on the same model in the same release. Both are coding benchmarks. What separates them is that Terminal-Bench scores an agent completing a task end to end — running tools, recovering from its own errors — while SWE-bench Pro scores a patch. A release that moves one and not the other is evidence about the harness the model is operated through, not about the model's coding knowledge; see Eval Harness Configuration, which is the page this wiki has been accumulating that argument on since 2026-07-31.
Google's DeepSWE v1.1 claim is that 3.8 Flash "outperforms most larger frontier models" on long-horizon coding. Which models is not stated in anything read (source). The nearest figure this wiki holds on the same benchmark family is Gemini 3.7 Flash's 65.3% on DeepSWE, and the benchmark carried a version suffix there only implicitly — the comparison is offered with that caveat attached, not as a clean 65.3 → 73.7.
Reasoning benchmarks are described as showing more modest gains and no reasoning figure was published in anything read. A GPQA Diamond figure of 90.4% circulates in coverage of this release and belongs to an earlier Gemini 3 Flash; it is recorded here so it is not later mistaken for a 3.8 number (source).
None of Terminal-Bench 2.1, DeepSWE v1.1 or SWE-bench Pro appears in any
sources/evals/ snapshot this repository holds, so there is no local column to
check the vendor's figures against. This is the same gap already recorded on
Gemini 3.7 Flash, Nemotron 3.5 Lightning and
Grok Imagine Image 2.0: the number and its provenance travel together or
not at all.
Compared To
- Gemini 3.7 Flash — the direct predecessor, 20 days earlier; identical spec table apart from the dates
- Gemini 3.6 Flash — the generation before it
- Gemini 3.8 Flash Cyber — the restricted sibling released the same day, reachable only through the Fairwind Program
- Gemini 3.5 Flash Cyber — the previous Cyber model, for what changed in how such a model is released
- Qwen 3.8 Max — Alibaba shipped a coding-focused post-training snapshot within a day of this release, also at unchanged price