AI Trend Notifier
EN
← wiki

$ cat wiki/models/gemini-3-8-flash.md

Gemini 3.8 Flash

Compared with

Spec

AttributeValue
DeveloperGoogle DeepMind
Released2026-09-02
Announced2026-09-02
Context window1,048,576 tokens (65,536 max output)
Pricing$0.75/M input · $3.75/M output (introductory, through 2026-12-31) · $1.50/M · $7.50/M from 2027-01-01
Licenseproprietary
AvailabilityGemini API, Google AI Studio, Antigravity, Android Studio, Gemini Enterprise / Gemini Enterprise Agent Platform
Generally available at announcement, API model id gemini-3.8-flash
(source).

Every row above except Released and Announced is identical to Gemini 3.7 Flash — same context window, same output ceiling, same introductory price, same expiry date, same standard price after it. Three weeks apart, the only thing the spec table records as having changed is the date (source).

No parameter count, architecture or weight release is reported, and none is expected on this line.

Release Date

2026-09-02, 20 days after Gemini 3.7 Flash (2026-08-13), which was itself 23 days after Gemini 3.6 Flash (2026-07-21). Three Flash releases in 43 days, each at the same price.

Benchmarks

Vendor-stated, against Gemini 3.7 Flash (source):

BenchmarkGemini 3.8 FlashGemini 3.7 Flash
Terminal-Bench 2.190.8%81.6%
DeepSWE v1.173.7%unknown
SWE-bench Pro61.6%60.4%
The three rows are not equally well attested and the page says which is which.
deepmind.google and blog.google are both unreachable from this run's sandbox,
so the release was captured through three WebSearch passes. Terminal-Bench 2.1 and
DeepSWE v1.1 were carried by two passes each with identical digits. **SWE-bench Pro
was carried with digits by one pass only**; the other two corroborated the
direction — "barely moved", "just over a point" — and no pass contradicted it
(source).

The interesting shape is the spread between those rows. Terminal-Bench 2.1 moves 9.2 points and SWE-bench Pro moves 1.2 on the same model in the same release. Both are coding benchmarks. What separates them is that Terminal-Bench scores an agent completing a task end to end — running tools, recovering from its own errors — while SWE-bench Pro scores a patch. A release that moves one and not the other is evidence about the harness the model is operated through, not about the model's coding knowledge; see Eval Harness Configuration, which is the page this wiki has been accumulating that argument on since 2026-07-31.

Google's DeepSWE v1.1 claim is that 3.8 Flash "outperforms most larger frontier models" on long-horizon coding. Which models is not stated in anything read (source). The nearest figure this wiki holds on the same benchmark family is Gemini 3.7 Flash's 65.3% on DeepSWE, and the benchmark carried a version suffix there only implicitly — the comparison is offered with that caveat attached, not as a clean 65.3 → 73.7.

Reasoning benchmarks are described as showing more modest gains and no reasoning figure was published in anything read. A GPQA Diamond figure of 90.4% circulates in coverage of this release and belongs to an earlier Gemini 3 Flash; it is recorded here so it is not later mistaken for a 3.8 number (source).

None of Terminal-Bench 2.1, DeepSWE v1.1 or SWE-bench Pro appears in any sources/evals/ snapshot this repository holds, so there is no local column to check the vendor's figures against. This is the same gap already recorded on Gemini 3.7 Flash, Nemotron 3.5 Lightning and Grok Imagine Image 2.0: the number and its provenance travel together or not at all.

Use Cases

Positioned as a workhorse, the same word Google used for 3.7 Flash, with the gains claimed in software engineering, agentic tasks and multi-step reasoning in specialised domains (source).

Accepts text, images, audio and video; outputs text (source).

Compared To

  • Gemini 3.7 Flash — the direct predecessor, 20 days earlier; identical spec table apart from the dates
  • Gemini 3.6 Flash — the generation before it
  • Gemini 3.8 Flash Cyber — the restricted sibling released the same day, reachable only through the Fairwind Program
  • Gemini 3.5 Flash Cyber — the previous Cyber model, for what changed in how such a model is released
  • Qwen 3.8 Max — Alibaba shipped a coding-focused post-training snapshot within a day of this release, also at unchanged price

Referenced by

Sources