Vals VoiceCodeBench (VoiceCodeBench)
We show this table for reference; we do not rank on it.
Tests whether speech-to-text models preserve exact structured values in English workplace audio.
Task success rate on VoiceCodeBench — September 24, 2026
We mirror the published task success rate view for VoiceCodeBench. GPT Live Transcribe leads the public snapshot at 67.67%, followed by Grok Voice Transcribe 2.0 (65.33%) and Ink 2 (62.00%). We do not use these results to rank models overall.
GPT Live Transcribe
OpenAI
Grok Voice Transcribe 2.0
xAI
Ink 2
Cartesia
20 modelsMultimodal & GroundedCurrentDisplay onlyUpdated September 24, 2026
Task success rate table (20 models)
ScoreHow VoiceCodeBench is shown here
BenchLM mirrors the public Vals AI VoiceCodeBench leaderboard captured from https://www.vals.ai/benchmarks/voice-code-bench and updated by Vals on September 24, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.
VoiceCodeBench is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.
Snapshot
The published VoiceCodeBench snapshot places GPT Live Transcribe first at 67.67%. The third row is 5.67 points behind. The broader top-10 range is 12.00 points, so the table still separates the published systems.
20 models have been evaluated on VoiceCodeBench. The benchmark falls in the Multimodal & Grounded category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. VoiceCodeBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About VoiceCodeBench
Year
2026
Tasks
300 workplace speech clips with exact structured values
Format
Task success rate
Difficulty
Structured-value speech transcription
The 300 human-recorded clips contain 1,482 audited target entities. Task success requires every target value in a clip to be correct; entity recovery and word error rate remain separate diagnostics. These STT results are display-only and do not enter text-model or voice-agent rankings.
Freshness and provenance
Version
VoiceCodeBench 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does VoiceCodeBench measure?
Tests whether speech-to-text models preserve exact structured values in English workplace audio.
Which model leads the published VoiceCodeBench snapshot?
GPT Live Transcribe currently leads the published VoiceCodeBench snapshot with 67.67% task success rate. BenchLM shows this benchmark for display only and does not use it in overall rankings.
How many models are evaluated on VoiceCodeBench?
The September 24, 2026 snapshot contains 20 AI models.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.