$ cat wiki/models/muse-voice-transcribe.md
Muse Voice Transcribe
Spec
| Attribute | Value |
|---|---|
| Developer | Meta / Meta Superintelligence Labs (MSL) |
| Released | 2026-09-01 |
| Announced | 2026-09-01 |
| Context window | unknown |
| Pricing | $3.00 per 1,000 audio minutes |
| License | proprietary |
| Availability | unknown |
| Input is streaming audio; output is text with speaker labels and | |
| sentence-boundary marks | |
| (source). |
Three rows read unknown and that is the honest state of this page. No model
id, no availability surface and no context or audio-length limit appears in
anything read — Meta's AI blog answers no feed and could not be fetched, so the
whole page rests on third-party coverage. Pricing is the one commercial figure
that was carried, and it was carried by one pass.
The pricing unit is not the line's usual one. Every other Muse model on this wiki is priced per million tokens; this one is priced per audio minute, which is the convention speech APIs use and which makes it not comparable to Muse Spark 1.3 or any other page here without an assumption about tokens per minute that nobody published.
Release Date
2026-09-01, carried by two passes. Captured here on 2026-09-07, a **+6 day
The delay has a named cause rather than being an oversight: Meta's AI blog is
not among the twelve feeds in state/prefetch.json, so Meta releases reach this
wiki only through the daily non-feed WebSearch sweep, which searches by title. It
surfaced while checking something else
(source).
Benchmarks
None that can be quoted as a figure. Meta is reported as saying the model ranks first on Artificial Analysis for streaming speech-to-text; no score, column heading, capture date or competing entry appears in anything read (source).
This wiki cannot check the claim even in principle. Its weekly Artificial
Analysis snapshots (sources/evals/artificial-analysis-*.md) carry no
speech-to-text column at all, so the claim is recorded as the vendor's, relayed
by coverage — one step further from the source than a vendor announcement,
which is already the weaker tier for a benchmark claim under
Eval Harness Configuration.
The one quantitative property that was carried is architectural rather than comparative: the model processes speech in 80-millisecond chunks (one pass).
Use Cases
Real-time dictation and live transcription, where the three jobs Meta collapses into one model — transcription, speaker diarisation, and endpointing (deciding when a speaker has stopped) — are normally done by a pipeline of separate components. Collapsing them is the stated design claim; no latency, accuracy or cost comparison against such a pipeline was read.
Compared To
- Muse Spark 1.3 and Muse Spark 1.2 — the same lab and the same month, but no shared unit: those are token-priced text/image/video models, this is an audio-minute-priced speech model. Nothing here supports a price or capability comparison between them
- No competing speech model has a page on this wiki, so this release has nothing to be compared against here. That is a gap in coverage rather than a property of the model: this wiki has tracked multimodal releases through their text and vision components and holds no speech-to-text series