$ cat wiki/models/gemini-3-5-flash.md
Gemini 3.5 Flash
Compared with
Spec
| Attribute | Value |
|---|---|
| Developer | Google DeepMind |
| Released | 2026-05-19 (GA at Google I/O 2026) |
| Announced | 2026-05-19 (Google I/O 2026) |
| Context window | 1,048,576 input / 65,536 output tokens |
| Pricing | $1.50 in / $9.00 out per 1M tokens ($0.15 cached) |
| License | proprietary (API-only; no weight release) |
| Availability | Gemini API, Google AI Studio, Google Antigravity, Gemini app, AI Mode in Google Search |
| Family | Gemini 3.5 |
| Modalities | Text + image + audio + video → text |
| Thinking | Dynamic thinking (on by default) |
| Knowledge cutoff | January 2026 |
| Speed | 4× faster than other frontier models (output tokens/sec) |
Benchmarks
| Benchmark | Score | Context |
|---|---|---|
| Terminal-Bench 2.1 | 76.2% | Agentic coding |
| MCP Atlas | 83.6% | MCP/tool use |
| CharXiv Reasoning | 84.2% | Chart reasoning |
| GPQA Diamond | 90.4% | PhD-level science |
| MMMU-Pro | 81.2% | Multimodal understanding |
| Outperforms Gemini 3.1 Pro across the full coding and agentic benchmark suite. |
Use Cases
Primary positioning: agentic and coding tasks at Flash speed and cost. The "strongest agentic and coding model the Flash series has ever shipped."
- Multi-step agent workflows (MCP Atlas 83.6%)
- Scientific reasoning (GPQA Diamond 90.4%)
- Coding tasks (Terminal-Bench 76.2%)
- Multimodal analysis
Availability
Gemini API, Google AI Studio, Google Antigravity, Gemini app, AI Mode in Google Search.
Compared To
- Gemini 3.1 Pro: Gemini 3.5 Flash surpasses it on coding/agentic/multimodal benchmarks at lower cost and 4× speed
- GPT-5.5 Instant: Both target fast, cost-efficient frontier tiers. Gemini 3.5 Flash has an explicit MCP Atlas benchmark; GPT-5.5 Instant lacks published agentic benchmarks
- Claude Opus 4.7: Anthropic flagship vs. Google's Flash tier; different price/performance tradeoffs
- Devstral 2: Open-weight coding specialist (72.2% SWE-bench) vs. Gemini 3.5 Flash's broader multimodal + agentic scope
Conflicting Reports
-
The published price disagrees with the catalogue's first-party endpoint, and the page's figure stands.
spec-checkrun 90 (2026-09-12) reports input $1.5 vs $0.75; output $9 vs $4.50, both 2.0× against Google (1차) (source).The Spec row is not changed. Per
CLAUDE.md, a page that cites a vendor announcement keeps the vendor's figure, and the cell above is faithful to the announcement this page cites. What was missing was the disclosure, which is what the schema asks for and what this entry supplies.What is not established: which figure is correct, and what the catalogue's price is a price for — no tier, context band, modality or billing unit appears in the check's output. The ratio is an exact small-integer multiple in both cells at once, and it is on six of the eight conflicting pages, which is the shape of a tier or unit mismatch rather than of six independent errors — recorded as a pattern and not adopted as an explanation (source).
This has been reported by a red Action on every run since 2026-08-29 — 25 consecutive — and had reached no page until 2026-09-13.
openrouter.aiis blocked from the daily run's sandbox, so the catalogue cannot be re-read here to settle it; that is a question for a run with egress.
Sources
- Google Blog launch
- llm-stats.com benchmarks
- Announced at Google I/O 2026 (2026-05-19) — see Google DeepMind
- Google — Gemini API pricing — spec figures verified 2026-07-27