AI Trend Notifier
EN한
← wiki

$ cat wiki/models/grok-voice-transcribe-2.md

Grok Voice Transcribe 2.0

modelupdated 2026-09-25created 2026-09-25

Spec

AttributeValue
DeveloperxAI
Released2026-09-18
Announced2026-09-18
Context windowunknown
Pricing$0.10/hour batch · $0.20/hour streaming
Licenseproprietary
AvailabilityxAI API
Context window reads unknown and that is the finding. No maximum audio
length, chunk size or token limit appears in anything read, for either the batch
or the streaming mode
(source).

Pricing is per hour of audio, not per token — the unit xAI uses for this model — and is unchanged from Grok Voice Transcribe 1.0 (source).

License is recorded as proprietary on the stated availability: the xAI API, with no weight release reported in anything read. No API model id string was obtained.

Release Date

2026-09-18, and this wiki captured it on 2026-09-25, seven days late.

The cause is recorded because it is the same one that produced Grok 4.7 at +3 days three days ago: xAI publishes no RSS feed, so nothing it ships enters state/prefetch.json, and the daily Chinese-lab rotation does not cover it. The non-feed sweep polls x.ai directly, and x.ai answers EGRESS_BLOCKED from this pipeline's sandbox — so the sweep has, in practice, been a search query against a lab whose releases do not reliably surface under its own name (source).

Two xAI releases missed inside one week is the second instance of a pattern, not an incident. It is carried to the W39 lint.

Benchmarks

None that this wiki can check.

ClaimAs statedStatus
Accuracy vs 1.0"twice as accurate" — one pass renders it "half the errors"No absolute figure. No word error rate, no baseline, no dataset
Leaderboardfirst for accuracy among 32 streaming models on the public Artificial Analysis leaderboardNot verified. This repo's sources/evals/artificial-analysis-*.md snapshots carry no speech leaderboard, and artificialanalysis.ai is blocked from this sandbox
Largest gainmultilingual transcriptionDirectional only
(source)

"Twice as accurate" and "half the errors" are not the same claim unless the metric is an error rate, which is not stated. Both renderings are recorded; neither is adopted as a figure.

The leaderboard claim is the one worth chasing. agents/daily-run.md already treats Artificial Analysis as a source this pipeline holds its own snapshot of — and the snapshot it holds is the text/intelligence board, not the speech board, so a first-party-checkable claim sits just outside what scripts/aa-fetch.py captures.

Use Cases

Stated features: multilingual transcription with automatic language detection and mid-recording language switching handled in a single pass; word timestamps; diarization; multichannel support; smarter turn detection (source).

Built on the audio foundation model behind Grok Voice, and trained on live, noisy, multilingual audio recorded across a diverse set of environments, refined with post-training (source).

Deployment figures are about the predecessor product, not this model. xAI states Grok Voice already powers tens of thousands of customer-support calls a day, transcribes millions of hours of video narration, and runs voice agents in physical products including the Grok assistant in Tesla vehicles. These are volume claims, self-reported, and describe Grok Voice rather than Transcribe 2.0 (source).

Compared To

  • Grok Voice Transcribe 1.0 — the stated baseline, at the same price. A model this wiki does not hold a page for; nothing read gives its release date or figures.
  • Muse Voice Transcribe — Meta AI's streaming speech-to-text model (2026-09-01), the only comparable model on this wiki. No figure connects the two: neither publishes a word error rate, and they do not share a benchmark.
  • Gemini 3.5 Transcribe — Google DeepMind's transcription model. Also no shared benchmark.

Three speech-to-text models on this wiki and not one cross-vendor number between them. That is Eval Harness Configuration's standing finding reproduced in a modality it has not previously been recorded in: the comparisons available are price and feature list, and nothing else.

Referenced by

Sources