$ cat wiki/models/gemini-3-8-flash-tts.md
Gemini 3.8 Flash TTS
Google DeepMind's expressive text-to-speech model, announced 2026-09-23 alongside Gemini 3.8 Flash-Lite TTS and rolling out the same day in the Gemini API and Google AI Studio. Positioned for deep creative direction and character design across gaming, immersive audiobooks, podcasts and interactive media (source).
Not read first-party. deepmind.google answers EGRESS_BLOCKED from this
run's sandbox; the title, URL and date come from Google DeepMind's own RSS feed
through state/prefetch.json and the body from two search passes.
Spec
| Attribute | Value |
|---|---|
| Developer | Google DeepMind |
| Released | 2026-09-23 |
| Announced | 2026-09-23 |
| Context window | 8,192 text input tokens |
| Pricing | unknown |
| License | unknown |
| Availability | Gemini API, Google AI Studio |
| Two rows need their reading stated | |
| (source): |
Context windowis a text input budget, not a conversational context. 8,192 tokens in, audio out. The row carries the documented number because the schema fixes the attribute name, and this is the only input-length figure published.Pricingisunknownand deliberately so. One pass offers $0.037 per minute of audio — for a prior Flash TTS model — and a per-token rate for the Gemini 3.8 Flash text model, which is a different model entirely. Neither is this model's price, and neither is adopted. This row is whatspec-checkwill read against the catalogue once a price exists.
Release Date
2026-09-23 (2 passes), announced and rolling out the same day.
Benchmarks
None. No benchmark, MOS score, listener-preference result or comparison figure of any kind appears in anything read. The entire quality claim is Google's own phrase — "its most expressive audio generation models yet" — with no measurement attached (source).
This is worth stating rather than leaving as a blank section. Speech synthesis has established public evaluations, and a release that publishes none is making a claim this wiki cannot check today and could not check later either, because nothing was pinned to compare against.
Use Cases
Stated by Google (2 passes) (source):
- Gaming — character design and voice direction
- Immersive audiobooks
- Podcasts
- Interactive media
The capability framing is prompt-directed emotional speech with "fine-grained control over tone, pacing, and expressive nuance" (1 pass), plus voice design and voice replication (1 pass). 130 languages are supported; the list is not published in anything read.
Every generated clip carries a SynthID watermark, described as imperceptible and embedded directly in the audio output so that AI-generated speech remains detectable (2 passes). That is the only safeguard named — and it is named beside voice replication, with no consent, likeness or misuse control appearing in anything read.
Compared To
- Gemini 3.8 Flash-Lite TTS — announced in the same post, same day, same API. The split is creative direction against volume: this model for character work, Flash-Lite for dubbing, bulk audio and voice agents. No figure separates them in anything read — not price, not latency, not quality — so the distinction is currently positioning and nothing else.
- Gemini 3.5 Transcribe — the other side of the audio line, speech in rather than speech out, and the nearest comparison this wiki holds.
- Gemini 3.8 Flash — shares the version number and nothing else; it is the text model and its price rows do not apply here.