$ cat wiki/models/gemini-3-8-live.md
Gemini 3.8 Live
Spec
| Attribute | Value |
|---|---|
| Developer | Google DeepMind |
| Released | 2026-09-15 |
| Announced | 2026-09-15 |
| Context window | unknown |
| Pricing | unknown |
| License | proprietary |
| Availability | Gemini API, Google AI Studio, Search Live; enterprise via Gemini Enterprise |
| Announced together with Gemini 3.8 Live Extended Thinking in a single | |
| post, "Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking", and | |
| described as rolling out "starting today" | |
| (source). |
Four rows read unknown and that is the finding, not an omission. No pass
returned a per-token price, a context window, a maximum output or an API model
id for either Live model. A pass that asked directly for "price per million
tokens" returned Gemini 3.8 Flash's figures instead and said so —
$0.75/M in, $3.75/M out, 1,048,576-token context — which belong to a
different model released thirteen days earlier and are recorded on that page,
not here (source).
License is recorded as proprietary on the stated availability: API, AI Studio
and Search Live, with no weight release reported in anything read.
Release Date
2026-09-15. The announcement is dated by DeepMind's own RSS entry —
Tue, 15 Sep 2026 17:05:57 +0000 — captured in state/prefetch.json and read
there rather than from the page, which is unreachable from this run's sandbox
(source).
Thirteen days after Gemini 3.8 Flash and Gemini 3.8 Flash Cyber (2026-09-02), and the third distinct Gemini 3.8 access envelope in a fortnight.
Benchmarks
None. No figure of any kind is reported for this model, on the Artificial Analysis Speech to Speech Quality Index or on anything else. Every published number in the release belongs to Gemini 3.8 Live Extended Thinking (source).
This is the same shape Gemini 3.8 Flash Cyber took on 2026-09-02 — two models announced together, one carrying the numbers — with the difference that the silent model there was the restricted one. Here it is the general-availability one, and no access restriction explains it.
Use Cases
Positioned as the scale-and-cost-efficiency option of the pair, built for fluid dialogue and visual grounding, against a sibling aimed at higher-complexity reasoning (1 pass) (source).
Capabilities stated for the pair, carried by one pass as a single list: detecting and switching between 97 languages mid-conversation, executing tool calls and API requests in the background while continuing to talk, and processing visual input in near real time (source).
The background-tool-call property is the one that touches Agents (LLM Agents): a voice model that can keep talking through a tool call is a different latency budget from one that cannot, and it is the capability a voice agent is built on rather than a conversational nicety. Nothing read quantifies it — no latency figure, no concurrency limit, no tool-call benchmark.
Compared To
- Gemini 3.8 Live Extended Thinking — announced the same day, same rollout surfaces; the sibling that carries every published figure
- Gemini 3.8 Flash — 2026-09-02, the 3.8 generation's general workhorse; its price and context window are published and are not this model's
- Gemini Omni 1.1 Flash — the other DeepMind line billed by media duration rather than by token
- GPT-Live-1 — OpenAI's live-audio line, the comparison the vendor chose for the sibling model
- Grok Voice Think Fast 2.0 — xAI's, likewise