Skip to main content
BenchLM

Vals VoiceCodeBench (VoiceCodeBench)

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

Tests whether speech-to-text models preserve exact structured values in English workplace audio.

Task success rate on VoiceCodeBench — September 24, 2026

We mirror the published task success rate view for VoiceCodeBench. GPT Live Transcribe leads the public snapshot at 67.67%, followed by Grok Voice Transcribe 2.0 (65.33%) and Ink 2 (62.00%). We do not use these results to rank models overall.

20 modelsMultimodal & GroundedCurrentDisplay onlyUpdated September 24, 2026

Task success rate table (20 models)

Score
1
GPT Live TranscribeOpenAI · Closed
67.67%
2
65.33%
3
Ink 2Cartesia
62.00%
4
59.67%
5
59.33%
6
59.00%
7
57.33%
9
Muse Voice TranscribeMeta · Closed
57.00%
10
55.67%
11
Chirp 3Google Cloud
55.33%
12
55.33%
13
54.33%
14
53.00%
18
Cohere Transcribe 03-2026Cohere · Open weight
44.67%
19
37.33%
20
30.33%

How VoiceCodeBench is shown here

BenchLM mirrors the public Vals AI VoiceCodeBench leaderboard captured from https://www.vals.ai/benchmarks/voice-code-bench and updated by Vals on September 24, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

VoiceCodeBench is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

20 Vals rows3 task viewspublic datasetTasks: Overall, CTEM, WERDisplay only

The published VoiceCodeBench snapshot places GPT Live Transcribe first at 67.67%. The third row is 5.67 points behind. The broader top-10 range is 12.00 points, so the table still separates the published systems.

20 models have been evaluated on VoiceCodeBench. The benchmark falls in the Multimodal & Grounded category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. VoiceCodeBench is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About VoiceCodeBench

Year

2026

Tasks

300 workplace speech clips with exact structured values

Format

Task success rate

Difficulty

Structured-value speech transcription

The 300 human-recorded clips contain 1,482 audited target entities. Task success requires every target value in a clip to be correct; entity recovery and word error rate remain separate diagnostics. These STT results are display-only and do not enter text-model or voice-agent rankings.

Freshness and provenance

Version

VoiceCodeBench 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does VoiceCodeBench measure?

Tests whether speech-to-text models preserve exact structured values in English workplace audio.

Which model leads the published VoiceCodeBench snapshot?

GPT Live Transcribe currently leads the published VoiceCodeBench snapshot with 67.67% task success rate. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on VoiceCodeBench?

The September 24, 2026 snapshot contains 20 AI models.

Last updated: September 24, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.