Skip to main content
BenchLM

Data as of September 30, 2026 · How the score is built

CurrentProprietaryNon-Reasoning

Released Sep 18, 2026N/A contextxAI model documentation

Grok Voice Transcribe 2.0

Decision readingGrok Voice Transcribe 2.0 is tracked, but not publicly ranked yet. Published speech or voice results from Grok Voice Transcribe 2.0 launch evaluation appear below. They do not produce a weighted text-model score.

Released Sep 18, 2026 — see all recent releases

Voice benchmark evidence

Grok Voice Transcribe 2.0 launch evaluation is an owner-defined voice evaluation. Its result stays separate from BenchLM's weighted text-model ranking.

Short-phrase multilingual WER
6.8%
Short-phrase WER, Grok Voice Transcribe 1.0
20.6%
Independent streaming-accuracy leaderboard rank
#1 of 32
Overall accuracy vs 1.0
2x
Batch price
$0.10 per hour
Streaming price
$0.20 per hour
Provider API ID
grok-voice-transcribe-2.0
Provider launch snapshot · xAI · 2026-09-18 · lower WER is better; the per-category charts publish no readable figures

Lineage

The sequence follows explicit supersedes links. Each score is estimated for that model; a relative can inform a sparse estimate but never sets a floor, so a newer release can score below an earlier one. Scores and prices remain blank when the corresponding public row or first-party rate is unavailable.

  1. Sep 18, 2026 · you are here

    Grok Voice Transcribe 2.0

    Not publicly ranked · Price not listed

Base entry

Radar

Grok Voice Transcribe 2.0 release history

Full release history

Radar confirmed these at the source. Use Grok Voice Transcribe 2.0 in your work? Explore Radar to follow supported changes and choose your alerts.

Radar

Spec sheet

Each documented value carries its source. Missing fields stay visible as not sourced or not published, rather than disappearing from the page.

API model ID
grok-voice-transcribe-2.0xAI model documentation
Context window
N/A
Maximum output
Not sourced yet
Knowledge cutoff
Not sourced yet
Input modalities
Not sourced yet
Output modalities
Not sourced yet
Parameters
Not disclosed by the provider
Availability
Released September 18, 2026 in the xAI Speech-to-Text API for batch and streaming transcription, with word-level timestamps and confidence scores, speaker diarization at no extra cost, multichannel transcription of up to 8 channels, key term biasing of up to 100 terms per request, text formatting, filler-word removal, and smart turn detection. Existing integrations pick up the accuracy change with no code edits. xAI says it will soon be the API default and that Grok Voice Transcribe 1.0 will be deprecated in the coming weeks; callers can pin grok-voice-transcribe-1.0 during the transition. BenchLM does not track the 1.0 model.
Cloud regions
Not tracked yet
Lifecycle
Current
API capabilities
Tool calling, structured outputs, and batch support are not tracked yet
Prompt caching
Not documented in the pricing recordxAI model documentation
Self-host
Weights are not published
Rate limits
Not tracked yet

How to read this profile

The visual layer above carries the decisions. These notes preserve the model, ranking, coverage, and family context behind the numbers.

We track Grok Voice Transcribe 2.0, but no weighted text-model benchmark result is published on the site yet. This page shows the metadata and separate protocol evidence we can verify now; a BenchLM score will appear only if compatible public evaluations land.

Grok Voice Transcribe 2.0 is a proprietary model with a N/A context window. No explicit reasoning mode is documented in this profile.

Released September 18, 2026 in the xAI Speech-to-Text API for batch and streaming transcription, with word-level timestamps and confidence scores, speaker diarization at no extra cost, multichannel transcription of up to 8 channels, key term biasing of up to 100 terms per request, text formatting, filler-word removal, and smart turn detection. Existing integrations pick up the accuracy change with no code edits. xAI says it will soon be the API default and that Grok Voice Transcribe 1.0 will be deprecated in the coming weeks; callers can pin grok-voice-transcribe-1.0 during the transition. BenchLM does not track the 1.0 model.

Grok Voice Transcribe 2.0 is a display-only speech-to-text model built on the audio foundation model behind Grok Voice. Transcription results remain separate from BenchLM’s weighted text-model ranking.

Speech or voice evidence appears above; text-model benchmark coverage remains empty.

Last updated September 30, 2026. Runtime fields remain blank until a sourced snapshot exists.

Questions

How does Grok Voice Transcribe 2.0 perform overall in AI benchmarks?

Grok Voice Transcribe 2.0 has published speech or voice results in Grok Voice Transcribe 2.0 launch evaluation. Those results remain separate from the weighted text-model ranking, so this profile does not assign a BenchLM score or overall rank. Use the linked evaluation's metric and run conditions to compare speech systems.

Does Grok Voice Transcribe 2.0 have full benchmark coverage on BenchLM?

No. Grok Voice Transcribe 2.0's speech or voice results appear in Grok Voice Transcribe 2.0 launch evaluation, but it has no published text-model benchmark rows. The voice evaluation does not fill the text-model coverage slots or produce an overall rank. Missing text coverage is not treated as a zero score.

What is the context window size of Grok Voice Transcribe 2.0?

Grok Voice Transcribe 2.0 has a reported context window of N/A in the exact-model catalog record. The value stays visible, but the profile marks its source link as unavailable instead of presenting it as directly documented. Maximum output length remains separate because providers often publish a different limit.

Watch Grok Voice Transcribe 2.0 in the weekly brief

Get one weekly email when material rank, price, availability, or benchmark evidence changes are worth revisiting.

Read a sample issue

Join 2,000+ readers.

Compare Grok Voice Transcribe 2.0 with every tracked model636 comparisons