Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes
Official provider results

Grok Voice Transcribe 2.0 launch evaluation

xAI’s September 18, 2026 launch post reports a rank on an independent streaming speech-to-text leaderboard, four internal word-error-rate sets drawn from production traffic, and one exact multilingual short-phrase pair against Grok Voice Transcribe 1.0.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
Four internal sets from production traffic — telephony (8 kHz customer-support calls, English), conversational (conversations with Grok, English), credentials (phone numbers, emails and addresses read aloud, English), and short phrases (voice-assistant utterances across 19 languages) — plus an independent public streaming speech-to-text leaderboard
Primary metric
Word error rate (lower is better)
Owner
xAI
Available evidence
The short-phrase pair and the pricing are stated in the post’s prose. The four internal charts and the multilingual comparison render client-side, so their per-category values are not published in a readable form.

What it measures

Voice experience
Is the exchange natural, robust, and responsive?

Available results

Short-phrase multilingual WER
6.8%
Down from 20.6%

Voice-assistant utterances across 19 languages, the set xAI calls its largest improvement. Grok Voice Transcribe 1.0 scored 20.6% on the same set. Short commands give the model little context to identify the language.

Independent streaming-accuracy leaderboard rank
#1 of 32
Provider-reported

xAI places the model first for accuracy among the 32 streaming models on a public third-party speech-to-text board. BenchLM does not track that board as a source, so the rank is recorded as an xAI claim rather than a measurement.

Telephony (8 kHz)
Not published
Provider-reported lead

Customer-support calls in English. xAI says the model leads every model it tested on this set and improves on Grok Voice Transcribe 1.0, but the chart carries no readable figures.

Conversational
Not published
Provider-reported gain

Conversations with Grok, English. xAI reports an improvement over Grok Voice Transcribe 1.0 without publishing the values.

Credentials
Not published
Provider-reported gain

Phone numbers, emails and addresses read aloud, English. xAI reports an improvement over Grok Voice Transcribe 1.0 without publishing the values.

Overall accuracy vs Grok Voice Transcribe 1.0
2x
Provider-reported

xAI summarises its real-world evaluations as twice as accurate as Grok Voice Transcribe 1.0 at the same price. The multilingual chart also compares against ElevenLabs Scribe v2 and Deepgram Nova-3, again without readable figures.

Price
$0.10 / $0.20 per hour
Unchanged from 1.0

Batch transcription $0.10 per hour of audio, streaming $0.20 per hour, with diarization, timestamps and key terms included.

Interpretation limit

The internal sets are xAI’s own production traffic, not a released dataset, so nothing here is independently reproducible. The four per-category charts and the multilingual bar chart are drawn in the browser and carry no readable figures, so BenchLM stores only the values the post states in prose and does not read numbers off the plots. The leaderboard rank is an xAI claim about a third-party board that BenchLM does not track as a source; it is not a BenchLM measurement. These transcription results stay separate from BenchLM’s weighted text-model ranking.