# Grok Voice Transcribe 2.0 launch evaluation: Voice Benchmark Profile

> xAI’s September 18, 2026 launch post reports a rank on an independent streaming speech-to-text leaderboard, four internal word-error-rate sets drawn from production traffic, and one exact multilingual short-phrase pair against Grok Voice Transcribe 1.0.

- Scope: Four internal sets from production traffic — telephony (8 kHz customer-support calls, English), conversational (conversations with Grok, English), credentials (phone numbers, emails and addresses read aloud, English), and short phrases (voice-assistant utterances across 19 languages) — plus an independent public streaming speech-to-text leaderboard
- Measurement lanes: Voice experience
- Primary metric: Word error rate (lower is better)
- Available evidence: The short-phrase pair and the pricing are stated in the post’s prose. The four internal charts and the multilingual comparison render client-side, so their per-category values are not published in a readable form.
- Owner: xAI
- Source snapshots refreshed: 2026-09-18

## Interpretation limit

The internal sets are xAI’s own production traffic, not a released dataset, so nothing here is independently reproducible. The four per-category charts and the multilingual bar chart are drawn in the browser and carry no readable figures, so BenchLM stores only the values the post states in prose and does not read numbers off the plots. The leaderboard rank is an xAI claim about a third-party board that BenchLM does not track as a source; it is not a BenchLM measurement. These transcription results stay separate from BenchLM’s weighted text-model ranking.

## Primary sources

- [Owner page](https://x.ai/news/grok-voice-transcribe-2)

Canonical page: https://benchlm.ai/voice-benchmarks/grok-voice-transcribe-2-launch-evaluation
