AI Trend Notifier
EN한
← wiki

$ cat wiki/models/gemini-3-7-flash.md

Gemini 3.7 Flash

Compared with

Spec

AttributeValue
DeveloperGoogle DeepMind
Released2026-08-13
Announced2026-08-13
Context window1,000,000 tokens (64,000 max output)
Pricing$0.75/M input · $3.75/M output (introductory, through 2026-12-31) · $1.50/M · $7.50/M from 2027-01-01
Licenseproprietary
AvailabilityGemini API, Google AI Studio, Android Studio, Google Antigravity, Gemini Enterprise Agent Platform, Gemini Enterprise app, Spark in the Gemini app (AI Pro and Ultra)
Generally available at announcement — a stable API model, not a preview
(source).

Knowledge cutoff March 2026, unchanged from its predecessor (source).

No parameter count, architecture or weight release is reported, and none is expected on this line.

Release Date

2026-08-13, 23 days after Gemini 3.6 Flash (2026-07-21).

Logan Kilpatrick, who leads product for Google AI Studio, described the gain as arriving in "roughly three weeks" and attributed it to algorithmic improvements across Google DeepMind teams rather than to scale (source). That is a claim about the source of the improvement and no evidence for it was published; it is recorded as the vendor's account.

Benchmarks

Vendor-stated, against Gemini 3.6 Flash (source):

BenchmarkGemini 3.7 FlashGemini 3.6 Flash
AutomationBench30.4%17.0%
FrontierCode43.6%34.4%
DeepSWE65.3%48.6%
LVBench (long video)85.4%unknown
GDM-MRCR v2 @128k97.0%unknown
GDM-MRCR v2 @1M62.5%unknown
WebDev Arena1588 Elounknown
On AutomationBench — enterprise workflow automation — coverage places it
ahead of GPT-5.6 Terra (23.6%) and **Claude Sonnet 5
(10.7%)** (source).
Whether Google ran those comparison figures itself is not stated in anything
read.

Not one of these benchmarks appears in any sources/evals/ snapshot this repository holds — not AutomationBench, FrontierCode, LVBench, GDM-MRCR v2 or WebDev Arena. This is the same gap recorded on Muse Glimmer, Nemotron 3.5 Lightning and Grok Imagine Image 2.0: there is no local column to check the vendor's figure against, so the number and its provenance travel together or not at all. See Eval Harness Configuration.

Note the shape of the headline: 30.4% is roughly double 17.0%, and it is also under a third of the tasks. Both readings are in the same figure and the release material carries only the first (source).

The DeepSWE figure of 65.3% is the one directly comparable to a number this wiki already holds: DeepSeek V4-Pro-0813 reports 62.7 on the same benchmark, vendor-run and with no published harness on either side.

Use Cases

Positioned as a workhorse: Google's own blog title is "our most intelligent workhorse model" (source). The benchmark selection — enterprise workflow automation, coding, long-context retrieval, long video — describes the intended workload more precisely than the positioning does.

Accepts text, images, audio and video (source).

Compared To

  • Gemini 3.8 Flash — the successor, 20 days later; identical spec table apart from the dates, with Terminal-Bench 2.1 up 9.2 points and SWE-bench Pro up 1.2
  • Gemini 3.6 Flash — the direct predecessor, 23 days earlier; same context window, same output ceiling, same standard price
  • Gemini 3.5 Flash — the generation before it
  • GPT-5.6 Sol (and Terra, Luna) — its Terra tier is the model named as the AutomationBench comparison
  • Claude Sonnet 5 — the other named comparison, at 10.7% on AutomationBench
  • DeepSeek V4-Pro-0813 — released the same day, and the nearest DeepSWE figure

Conflicting Reports

  • The published price disagrees with the catalogue's first-party endpoint, and the page's figure stands. spec-check run 90 (2026-09-12) reports input $0.75 vs $0.38; output $3.75 vs $1.88, both ~2.0× against Google (1차) (source).

    The Spec row is not changed. Per CLAUDE.md, a page that cites a vendor announcement keeps the vendor's figure, and the cell above is faithful to the announcement this page cites. What was missing was the disclosure, which is what the schema asks for and what this entry supplies.

    What is not established: which figure is correct, and what the catalogue's price is a price for — no tier, context band, modality or billing unit appears in the check's output. The ratio is an exact small-integer multiple in both cells at once, and it is on six of the eight conflicting pages, which is the shape of a tier or unit mismatch rather than of six independent errors — recorded as a pattern and not adopted as an explanation (source).

    This has been reported by a red Action on every run since 2026-08-29 — 25 consecutive — and had reached no page until 2026-09-13. openrouter.ai is blocked from the daily run's sandbox, so the catalogue cannot be re-read here to settle it; that is a question for a run with egress.

Referenced by

Sources