$ cat wiki/models/grok-voice-transcribe-2.md
Grok Voice Transcribe 2.0
Spec
| Attribute | Value |
|---|---|
| Developer | xAI |
| Released | 2026-09-18 |
| Announced | 2026-09-18 |
| Context window | unknown |
| Pricing | $0.10/hour batch · $0.20/hour streaming |
| License | proprietary |
| Availability | xAI API |
Context window reads unknown and that is the finding. No maximum audio | |
| length, chunk size or token limit appears in anything read, for either the batch | |
| or the streaming mode | |
| (source). |
Pricing is per hour of audio, not per token — the unit xAI uses for this
model — and is unchanged from Grok Voice Transcribe 1.0
(source).
License is recorded as proprietary on the stated availability: the xAI API,
with no weight release reported in anything read. No API model id string was
obtained.
Release Date
2026-09-18, and this wiki captured it on 2026-09-25, seven days late.
The cause is recorded because it is the same one that produced
Grok 4.7 at +3 days three days ago: xAI publishes no RSS feed, so
nothing it ships enters state/prefetch.json, and the daily Chinese-lab rotation
does not cover it. The non-feed sweep polls x.ai directly, and x.ai answers
EGRESS_BLOCKED from this pipeline's sandbox — so the sweep has, in practice,
been a search query against a lab whose releases do not reliably surface under
its own name
(source).
Two xAI releases missed inside one week is the second instance of a pattern, not an incident. It is carried to the W39 lint.
Benchmarks
None that this wiki can check.
| Claim | As stated | Status |
|---|---|---|
| Accuracy vs 1.0 | "twice as accurate" — one pass renders it "half the errors" | No absolute figure. No word error rate, no baseline, no dataset |
| Leaderboard | first for accuracy among 32 streaming models on the public Artificial Analysis leaderboard | Not verified. This repo's sources/evals/artificial-analysis-*.md snapshots carry no speech leaderboard, and artificialanalysis.ai is blocked from this sandbox |
| Largest gain | multilingual transcription | Directional only |
| (source) |
"Twice as accurate" and "half the errors" are not the same claim unless the metric is an error rate, which is not stated. Both renderings are recorded; neither is adopted as a figure.
The leaderboard claim is the one worth chasing. agents/daily-run.md already
treats Artificial Analysis as a source this pipeline holds its own snapshot of —
and the snapshot it holds is the text/intelligence board, not the speech
board, so a first-party-checkable claim sits just outside what
scripts/aa-fetch.py captures.
Use Cases
Stated features: multilingual transcription with automatic language detection and mid-recording language switching handled in a single pass; word timestamps; diarization; multichannel support; smarter turn detection (source).
Built on the audio foundation model behind Grok Voice, and trained on live, noisy, multilingual audio recorded across a diverse set of environments, refined with post-training (source).
Deployment figures are about the predecessor product, not this model. xAI states Grok Voice already powers tens of thousands of customer-support calls a day, transcribes millions of hours of video narration, and runs voice agents in physical products including the Grok assistant in Tesla vehicles. These are volume claims, self-reported, and describe Grok Voice rather than Transcribe 2.0 (source).
Compared To
- Grok Voice Transcribe 1.0 — the stated baseline, at the same price. A model this wiki does not hold a page for; nothing read gives its release date or figures.
- Muse Voice Transcribe — Meta AI's streaming speech-to-text model (2026-09-01), the only comparable model on this wiki. No figure connects the two: neither publishes a word error rate, and they do not share a benchmark.
- Gemini 3.5 Transcribe — Google DeepMind's transcription model. Also no shared benchmark.
Three speech-to-text models on this wiki and not one cross-vendor number between them. That is Eval Harness Configuration's standing finding reproduced in a modality it has not previously been recorded in: the comparisons available are price and feature list, and nothing else.
Sources
- SpaceXAI — Introducing Grok Voice Transcribe 2.0 → snapshot — EGRESS_BLOCKED, not read first-party
- MarkTechPost — SpaceXAI Releases Grok Voice Transcribe 2.0
- Unite.AI — xAI Releases Grok Voice Transcribe 2.0 Speech-to-Text Model
- AlphaSignal — xAI Ships Grok Voice Transcribe 2.0 With Half the Errors at Same Price