$ cat wiki/models/gemini-3-5-flash-lite.md
Gemini 3.5 Flash-Lite
Compared with
Spec
| Attribute | Value |
|---|---|
| Developer | Google DeepMind |
| Released | 2026-07-21 |
| Announced | 2026-07-21 |
| Context window | 1,048,576 tokens (1M) |
| Pricing | $0.30/M input · $2.50/M output |
| License | proprietary |
| Availability | Gemini API, Google AI Studio, Vertex AI |
Benchmarks
| Benchmark | Gemini 3.5 Flash-Lite | Gemini 3.1 Flash-Lite (predecessor) |
|---|---|---|
| Terminal-Bench 2.1 | 54% | 31% |
| GDM-MRCR v2 (long-context) | 72.2% | 60.1% |
| GDPval-AA v2 | 1140 | 642 |
| AA Intelligence Index | +11 pts over predecessor | baseline |
| Time per task | ~halved vs. predecessor | baseline |
| Output speed | 350.2 t/s | unknown |
Use Cases
- High-throughput, low-latency agentic pipelines (agentic search, document processing)
- Cost-sensitive applications requiring millions of calls per day
- Fast classification, routing, and summarization at scale
Key Differentiators
- Cheapest production Gemini model: $0.30/$2.50 per 1M tokens
- Speed: 350.2 t/s — significantly above the category median (~107 t/s)
- Configurable thinking levels: supports fast and deep reasoning modes
- Positioned as the high-throughput tier complementing Gemini 3.6 Flash
Related
- Google DeepMind — developer
- Gemini 3.6 Flash — higher-performance sibling, released same day
- Gemini 3.5 Flash Cyber — security-specialized sibling, released same day
- Agents (LLM Agents) — primary deployment context
Conflicting Reports
-
The published price disagrees with the catalogue's first-party endpoint, and the page's figure stands.
spec-checkrun 90 (2026-09-12) reports input $0.3 vs $0.15; output $2.5 vs $1.25, both 2.0× against Google (1차) (source).The Spec row is not changed. Per
CLAUDE.md, a page that cites a vendor announcement keeps the vendor's figure, and the cell above is faithful to the announcement this page cites. What was missing was the disclosure, which is what the schema asks for and what this entry supplies.What is not established: which figure is correct, and what the catalogue's price is a price for — no tier, context band, modality or billing unit appears in the check's output. The ratio is an exact small-integer multiple in both cells at once, and it is on six of the eight conflicting pages, which is the shape of a tier or unit mismatch rather than of six independent errors — recorded as a pattern and not adopted as an explanation (source).
This has been reported by a red Action on every run since 2026-08-29 — 25 consecutive — and had reached no page until 2026-09-13.
openrouter.aiis blocked from the daily run's sandbox, so the catalogue cannot be re-read here to settle it; that is a question for a run with egress.
Sources
- Google Blog (July 21, 2026): https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-6-flash-3-5-flash-lite-3-5-flash-cyber/
- DeepMind Flash-Lite page: https://deepmind.google/models/gemini/flash-lite/