# Vals VoiceCodeBench (VoiceCodeBench)

> Tests whether speech-to-text models preserve exact structured values in English workplace audio.

Canonical page: https://benchlm.ai/benchmarks/voice-code-bench

- Category: [Multimodal & Grounded](/multimodal-grounded)
- Last updated: September 24, 2026

## About VoiceCodeBench

- Year: 2026
- Tasks: 300 workplace speech clips with exact structured values
- Format: Task success rate
- Difficulty: Structured-value speech transcription
- Paper: [Vals VoiceCodeBench](https://www.vals.ai/benchmarks/voice-code-bench)

The 300 human-recorded clips contain 1,482 audited target entities. Task success requires every target value in a clip to be correct; entity recovery and word error rate remain separate diagnostics. These STT results are display-only and do not enter text-model or voice-agent rankings.

VoiceCodeBench is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (20 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GPT Live Transcribe](/models/gpt-live-transcribe) | OpenAI | 67.67% |
| 2 | [Grok Voice Transcribe 2.0](/models/grok-voice-transcribe-2) | xAI | 65.33% |
| 3 | [Ink 2](https://www.vals.ai/models/cartesia_ink-2) | Cartesia | 62.00% |
| 4 | [GPT-4o Transcribe](https://www.vals.ai/models/openai_gpt-4o-transcribe) | OpenAI | 59.67% |
| 5 | [Nova 3 Streaming](https://www.vals.ai/models/deepgram_nova-3-streaming) | Deepgram | 59.33% |
| 6 | [Resonant 1](https://www.vals.ai/models/reson8_resonant-1) | Reson8 | 59.00% |
| 7 | [Inworld Stt 1](https://www.vals.ai/models/inworld_inworld-stt-1) | Inworld | 57.33% |
| 8 | [Whisper Large V3 Turbo](https://www.vals.ai/models/groq_whisper-large-v3-turbo) | Groq | 57.00% |
| 9 | [Muse Voice Transcribe](/models/muse-voice-transcribe) | Meta | 57.00% |
| 10 | [Flux General En](https://www.vals.ai/models/deepgram_flux-general-en) | Deepgram | 55.67% |
| 11 | [Chirp 3](https://www.vals.ai/models/google_cloud_chirp_3) | Google Cloud | 55.33% |
| 12 | [Scribe V2 Realtime](https://www.vals.ai/models/elevenlabs_scribe_v2_realtime) | Elevenlabs | 55.33% |
| 13 | [Whisper Large V3](https://www.vals.ai/models/groq_whisper-large-v3) | Groq | 54.33% |
| 14 | [Voxtral Mini 2602](https://www.vals.ai/models/mistral_voxtral-mini-2602) | Mistral | 53.00% |
| 15 | [Grok Voice Transcribe 1.0](https://www.vals.ai/models/grok_grok-voice-transcribe-1.0) | xAI | 52.00% |
| 16 | [GPT-4o Mini Transcribe Streaming Response](https://www.vals.ai/models/openai_gpt-4o-mini-transcribe-streaming-response) | OpenAI | 51.67% |
| 17 | [Voxtral Mini Transcribe Realtime 2602](https://www.vals.ai/models/mistral_voxtral-mini-transcribe-realtime-2602) | Mistral | 46.00% |
| 18 | [Cohere Transcribe 03-2026](/models/cohere-transcribe-03-2026) | Cohere | 44.67% |
| 19 | [Universal Language Model](https://www.vals.ai/models/azure_speech_universal-language-model) | Azure_Speech | 37.33% |
| 20 | [Universal 3.5 Pro](https://www.vals.ai/models/assemblyai_universal-3-5-pro) | Assemblyai | 30.33% |

## FAQ

### What does VoiceCodeBench measure?

Tests whether speech-to-text models preserve exact structured values in English workplace audio.

### Which model leads the published VoiceCodeBench snapshot?

GPT Live Transcribe currently leads the published VoiceCodeBench snapshot with a score of 67.67%.

### How many models are evaluated on VoiceCodeBench?

The September 24, 2026 contains 20 AI models.

### Does VoiceCodeBench affect BenchLM's overall score?

Not directly. VoiceCodeBench is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
