# Whisper large-v3-turbo Benchmark Scores & Performance

> Whisper large-v3-turbo has Fermion speech evaluation results. Voice evidence is display-only and does not produce a text-model score.

Canonical page: https://benchlm.ai/models/whisper-large-v3-turbo

Last updated: September 30, 2026

## Model Details

| Property | Value |
|----------|-------|
| Creator | OpenAI |
| Source Type | Open Weight |
| Reasoning Type | Non-Reasoning |
| Context Window | N/A |
| Official model card | [OpenAI model documentation](https://huggingface.co/openai/whisper-large-v3-turbo) |
| Overall Score | Not scored (voice evidence only) |
| Overall Rank | Unranked |

## Family & Coverage

- Family: Whisper large-v3-turbo
- Variant: base
- Benchmarks covered: 0 of 486
- Coverage note: Speech and voice results appear below; no weighted text-model benchmark rows are stored.

## Speech and voice evidence

[Phonon-2 launch evaluation](/voice-benchmarks/phonon-2-launch-evaluation)

Fermion ran Phonon-2, Phonon-1, and Parakeet Redux with the Open ASR Leaderboard code on full test sets. The other accuracy rows are leaderboard figures quoted by Fermion, without independent re-verification here. Some model cards report different runs; those values are not mixed into this comparison. Voxtral's download size is an estimate at 16 bits. Throughput includes log-mel processing and excludes model load; batch-128 GPU results are separate from single-stream results. The Mac runtime comparison uses 20 dictations totaling 797 seconds at each runtime's defaults. Download size does not measure runtime memory. These results do not enter text-model or voice-agent rankings.

| Metric | Result |
| --- | --- |
| Seven-set mean WER | 6.58% |
| LibriSpeech clean WER | 2.13% |
| LibriSpeech other WER | 3.71% |
| AMI WER | 13.88% |
| Earnings-22 WER | 8.09% |
| GigaSpeech WER | 8.47% |
| SPGISpeech WER | 2.79% |
| VoxPopuli WER | 7.02% |
| Download | 1,618 MB |
| whisper.cpp, large-v3-turbo (Metal) | 17.0x realtime on M5 MacBook Air |

Leaderboard row quoted by Fermion; Fermion launch table retrieved September 30, 2026. Lower WER is better; throughput excludes model load.

- [Evaluation source](https://www.fermionresearch.com/research/phonon-2/)

Open multilingual speech recognition weights under MIT. The official model card documents a faster version of Whisper large-v3 with four decoder layers and inference through Transformers.

## Other OpenAI Models

- [GPT-6 Astra](/models/gpt-6-astra) - Score: 88.23
- [GPT-6 Sol](/models/gpt-6-sol) - Score: 78.58
- [GPT-5.6 Sol](/models/gpt-5-6-sol) - Score: 78.19
- [GPT-5.6 Terra](/models/gpt-5-6-terra) - Score: 72.41
- [GPT-5.5 Pro](/models/gpt-5-5-pro) - Score: 72.19
- [GPT-5.4 Pro](/models/gpt-5-4-pro) - Score: 70.68
- [GPT-5.5](/models/gpt-5-5) - Score: 68.77
- [GPT-5.4](/models/gpt-5-4) - Score: 68.23
- [GPT-6.1 Sol](/models/gpt-6-1-sol) - Score: 66.87
- [GPT-5.2 Pro](/models/gpt-5-2-pro) - Score: 66.75
- [GPT-6 Luna](/models/gpt-6-luna) - Score: 66.08
- [GPT-5.6 Luna](/models/gpt-5-6-luna) - Score: 65.39
- [GPT-5.3 Codex](/models/gpt-5-3-codex) - Score: 61.82
- [GPT-5.2](/models/gpt-5-2) - Score: 61.18
- [GPT-5.1](/models/gpt-5-1) - Score: 58.18
- [GPT-5.4 mini](/models/gpt-5-4-mini) - Score: 54.86
- [GPT-5.4 nano](/models/gpt-5-4-nano) - Score: 50.71
- [GPT-5.2-Codex](/models/gpt-5-2-codex) - Score: 49.87
- [o3-pro](/models/o3-pro) - Score: 47.14
- [o3](/models/o3) - Score: 46.96
- [GPT-5 (medium)](/models/gpt-5-medium) - Score: 44.65
- [GPT-5.1-Codex](/models/gpt-5-1-codex) - Score: 44.39
- [GPT-5 mini](/models/gpt-5-mini) - Score: 43.99
- [o3-mini](/models/o3-mini) - Score: 40.12
- [GPT-4.1](/models/gpt-4-1) - Score: 39.75
- [GPT-OSS 120B](/models/gpt-oss-120b) - Score: 37.67
- [o1](/models/o1) - Score: 37.49
- [o1-preview](/models/o1-preview) - Score: 36.26
- [o1-pro](/models/o1-pro) - Score: 35.71
- [GPT-5 nano](/models/gpt-5-nano) - Score: 34.97
- [GPT-OSS 20B](/models/gpt-oss-20b) - Score: 33.58
- [GPT-4o](/models/gpt-4o) - Score: 30.75
- [GPT-4.1 mini](/models/gpt-4-1-mini) - Score: 29.1
- [GPT-4o mini](/models/gpt-4o-mini) - Score: 26.88
- [GPT-4.1 nano](/models/gpt-4-1-nano) - Score: 24.9
- [GPT-4 Turbo](/models/gpt-4-turbo) - Score: 21.65
- [GPT-5.1-Codex-Max](/models/gpt-5-1-codex-max) - Score: not computed
- [o4-mini (high)](/models/o4-mini-high) - Score: not computed
- [GPT-5.2 Instant](/models/gpt-5-2-instant) - Score: not computed
- [GPT-5.3 Instant](/models/gpt-5-3-instant) - Score: not computed
- [GPT-5 (high)](/models/gpt-5-high) - Score: not computed
- [GPT-5.3-Codex-Spark](/models/gpt-5-3-codex-spark) - Score: not computed
- [GPT-Live-1](/models/gpt-live-1) - Score: not computed
- [GPT-5.6 Cyber](/models/gpt-5-6-cyber) - Score: not computed
- [GPT Live Transcribe](/models/gpt-live-transcribe) - Score: not computed
- [GPT Transcribe](/models/gpt-transcribe) - Score: not computed
- [GPT Realtime 2.1](/models/gpt-realtime-2-1) - Score: not computed
- [GPT Realtime 2.1 Mini](/models/gpt-realtime-2-1-mini) - Score: not computed
- [GPT Audio 1.5](/models/gpt-audio-1-5) - Score: not computed
- [GPT Realtime 1.5](/models/gpt-realtime-1-5) - Score: not computed
- [GPT-4o mini TTS](/models/gpt-4o-mini-tts) - Score: not computed
- [GPT Realtime](/models/gpt-realtime) - Score: not computed
- [GPT Realtime 2](/models/gpt-realtime-2) - Score: not computed
- [GPT Realtime mini](/models/gpt-realtime-mini) - Score: not computed
- [GPT-4o Audio](/models/gpt-4o-audio) - Score: not computed
- [GPT-4o mini Audio](/models/gpt-4o-mini-audio) - Score: not computed
