Muse Voice Transcribe launch evaluation
Meta’s launch evaluation reports streaming transcription accuracy, speaker diarization, and real-time audio-perception capabilities for Muse Voice Transcribe.
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Benchmark profile
- Scope
- Final-transcription streaming WER; average diarization error rate across AMI-IHM, AMI-SDM, and VoxConverse; multilingual and long-context capability disclosures
- Primary metric
- Word error rate and diarization error rate (lower is better)
- Owner
- Meta Superintelligence Labs
- Available evidence
- Exact figures published in the official launch post
What it measures
Available results
- Final-transcription streaming WER
- 3.1%
- Average diarization error rate
- 17.5%
- Time to final transcription
- About 0.16s
The other systems in Meta’s September 1 comparison range from 3.4% to 4.0%; lower is better.
Average DER across AMI-IHM, AMI-SDM, and VoxConverse; the other systems range from 21.1% to 28.6%.
The launch chart places Muse below the previous speed–accuracy frontier; the plotted value is approximate.