$ cat wiki/models/gemini-3-8-live-extended-thinking.md
Gemini 3.8 Live Extended Thinking
Spec
| Attribute | Value |
|---|---|
| Developer | Google DeepMind |
| Released | 2026-09-15 |
| Announced | 2026-09-15 |
| Context window | unknown |
| Pricing | $3.50 per hour of input audio |
| License | proprietary |
| Availability | Gemini API, Google AI Studio, Search Live; enterprise via Gemini Enterprise |
| Announced together with Gemini 3.8 Live in a single post, | |
| "Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking", and described as | |
| receiving "the same developer and enterprise rollout" | |
| (source). |
The Pricing cell is a rate per hour of audio, not a per-token price, and the
two are not interchangeable. It is the only cost figure published for either
model, it rests on one search pass, and no per-token rate, context window or
API model id appears in anything read. Written in the row rather than as
unknown because it is a real published price for a real billing unit; the
per-token price this repo's spec table normally carries does not exist in
anything read (source).
That matters beyond this page: scripts/spec-check.py reads Pricing back
against a per-token catalogue, so an audio-hour rate is a cell it cannot check.
This is the same gap Gemini Omni 1.1 Flash carries with its
per-second video pricing.
Release Date
2026-09-15, dated from DeepMind's own RSS entry
(Tue, 15 Sep 2026 17:05:57 +0000) captured in state/prefetch.json, since the
announcement page is unreachable from this run's sandbox
(source).
Benchmarks
Artificial Analysis Speech to Speech Quality Index — two passes, identical digits, same three rows, quoted with the column heading they were published under (source):
| Model | Speech to Speech Quality Index |
|---|---|
| Gemini 3.8 Live Extended Thinking | 82.6% |
| GPT-Live-1 "GPT-Live-1 Astra (Medium)" | 81.5% |
| Grok Voice Think Fast 2.0 "(High)" | 81.3% |
| Reported as taking "the top spot" on that index. |
The margin is 1.1 points over the second row and 1.3 over the third, and no confidence interval is published. Two of the three rows carry a reasoning-effort qualifier in parentheses — "Medium", "High" — and the winning row carries none, so the three entries are not stated to be at matched effort. This is the distinction Eval Harness Configuration exists for, and it is unresolved here.
This figure did not come from an sources/evals/artificial-analysis-*
snapshot. It is a vendor-and-coverage figure quoting Artificial Analysis's
column. The newest AA snapshot this repo holds is 2026-09-13 and carries no
speech-to-speech column, so there is no local column to check it against —
the same gap recorded on Gemini 3.8 Flash and
Gemini 3.7 Flash. AA snapshots are captured Sundays only, so the
earliest a local read could corroborate this is 2026-09-20.
Cost against the same two rivals (1 pass)
| Model | Cost per hour of input audio |
|---|---|
| Gemini 3.8 Live Extended Thinking | $3.50 |
| Grok Voice Think Fast 2.0 | $4.80 |
| GPT-Live-1 | $5.83 |
| Presented as undercutting both while topping them on quality — 27% below the | |
| nearer of the two. The quality table and the cost table name the same three | |
| models in the same order, which is the vendor's chosen frame and is recorded as | |
| such. |
Use Cases
Aimed at higher-complexity tasks that need more reasoning without breaking the flow of conversation (2 passes), against a sibling positioned for scale and cost efficiency (source).
Shares the pair's stated capabilities: detecting and switching between 97 languages mid-conversation, executing tool calls and API requests in the background while continuing to talk, and processing visual input in near real time (1 pass) (source).
"Extended thinking" in a live-audio setting is Test-Time Compute (Inference-Time Compute Scaling) placed under a latency constraint the text setting does not have: the reasoning budget has to fit inside a conversational turn. Nothing read states a thinking budget, a turn latency, or how the budget is controlled, so what the name denotes here is not established.
Compared To
- Gemini 3.8 Live — announced the same day, same surfaces; the cheaper half of the pair and the one with no published figure at all
- GPT-Live-1 — the OpenAI line the vendor benchmarks against, at 81.5% and $5.83/hour in these tables
- Grok Voice Think Fast 2.0 — the xAI line, at 81.3% and $4.80/hour
- Gemini 3.5 Transcribe — the same lab's speech line on the recognition side rather than the conversational one
- Gemini 3.8 Flash — the text workhorse of the same generation, whose price and context window are published