$ cat wiki/models/gemini-3-8-flash-lite-tts.md
Gemini 3.8 Flash-Lite TTS
Google DeepMind's high-volume text-to-speech model, announced 2026-09-23 alongside Gemini 3.8 Flash TTS and rolling out the same day in the Gemini API and Google AI Studio. Positioned for high-volume, cost-efficient use — dubbing, audio content creation and voice agents (source).
Not read first-party. deepmind.google answers EGRESS_BLOCKED from this
run's sandbox; the title, URL and date come from Google DeepMind's own RSS feed
through state/prefetch.json and the body from two search passes.
Spec
| Attribute | Value |
|---|---|
| Developer | Google DeepMind |
| Released | 2026-09-23 |
| Announced | 2026-09-23 |
| Context window | 8,192 text input tokens |
| Pricing | unknown |
| License | unknown |
| Availability | Gemini API, Google AI Studio |
Pricing is unknown, and on this model that is the row that matters. The | |
| entire stated reason for this model to exist separately from | |
| Gemini 3.8 Flash TTS is that it is cheaper at volume, and **no price | |
| for either model appears in anything read** — so the one claim distinguishing | |
| them is the one with no number under it. One pass records a prior Flash TTS | |
| model at approximately $0.037 per minute of audio; that is a different model | |
| and is not adopted here | |
| (source). |
The documented specification is otherwise identical to the Flash TTS model in everything published: same 8,192-token text input, same audio output, same 130 supported languages, same voice design and replication features, same SynthID watermarking (1 pass for the documentation set).
Release Date
2026-09-23 (2 passes), announced and rolling out the same day.
Benchmarks
None. No benchmark, MOS score, listener-preference result, latency figure or throughput figure appears in anything read (source).
For a model whose stated purpose is cost-efficient volume, latency and throughput are the figures that would establish the claim, and neither is published. The gap is recorded rather than filled.
Use Cases
Stated by Google (2 passes) (source):
- Dubbing
- Audio content creation at volume
- Voice agents
Voice agents is the entry that connects this model to the rest of the wiki. It is the first Google speech model on this wiki positioned at the agent stack rather than at media production, which puts it beside Agents (LLM Agents) rather than only beside the audio line.
Every generated clip carries a SynthID watermark (2 passes) — the same and only named safeguard as on the Flash model, and the same silence on consent and likeness controls beside a stated voice replication capability.
Compared To
- Gemini 3.8 Flash TTS — announced together, same day, same API, same published specification. The split is stated positioning with no figure behind it: creative direction and character design there, volume and cost here. Until a price or a latency number exists for either, this wiki holds two model pages that differ only in what Google says they are for.
- Gemini 3.5 Transcribe — speech in rather than speech out, the nearest audio comparison held here.