# Phonon-2 launch evaluation: Voice Benchmark Profile

> Phonon-2 has a reported 5.21% mean word error rate across seven English test sets. Fermion's launch report compares eight speech models and measures throughput across hardware and Mac runtimes.

- Scope: LibriSpeech test-clean and test-other, AMI, Earnings-22, GigaSpeech, SPGISpeech, and VoxPopuli; Phonon-2 hardware throughput; six runtime configurations on one M5 MacBook Air
- Measurement lanes: Voice experience, Latency
- Primary metric: Word error rate in percent (lower is better); audio seconds per wall-clock second (higher is better)
- Available evidence: Exact tables in the launch report, retrieved September 30, 2026; its accuracy table is also on the official model card
- Owner: Fermion Research
- Source snapshots refreshed: 2026-09-29

- Provider report retrieved: 2026-09-30

## English transcription accuracy

Word error rate (%), lower is better. Published averages are retained as reported. Download sizes describe weight files.

| Model | Download | LS clean | LS other | AMI | Earnings-22 | GigaSpeech | SPGISpeech | VoxPopuli | Mean WER | Evidence |
| --- | --- | --- | --- | --- | --- | --- | --- | --- | --- | --- |
| [Phonon-2](/models/phonon-2) | 164 MB | 1.72 | 3.92 | 9.37 | 6.96 | 8.35 | 3.70 | 2.46 | 5.21 | Fermion-run |
| [Parakeet TDT 0.6B v3](/models/parakeet-tdt-0-6b-v3) | 2,508 MB | 1.52 | 3.13 | 9.42 | 5.85 | 7.99 | 3.63 | 3.19 | 4.96 | Leaderboard row quoted by Fermion |
| [Parakeet Redux](/models/parakeet-redux) | 178 MB | 1.94 | 4.35 | 9.16 | 7.90 | 8.62 | 4.01 | 3.87 | 5.69 | Fermion-run |
| [Phonon-1](/models/phonon-1) | 415 MB | 2.11 | 5.03 | 10.31 | 12.34 | 8.73 | 3.67 | 3.73 | 6.56 | Fermion-run |
| [Canary 180M Flash](/models/canary-180m-flash) | 737 MB | 1.52 | 3.42 | 12.09 | 8.33 | 8.87 | 2.04 | 3.57 | 5.69 | Leaderboard row quoted by Fermion |
| [Voxtral Mini 4B Realtime 2602](/models/voxtral-mini-4b-realtime-2602) | About 8,000 MB (estimated) | 1.62 | 4.94 | 13.34 | 9.31 | 8.80 | 2.23 | 2.60 | 6.12 | Leaderboard row quoted by Fermion |
| [Whisper large-v3-turbo](/models/whisper-large-v3-turbo) | 1,618 MB | 2.13 | 3.71 | 13.88 | 8.09 | 8.47 | 2.79 | 7.02 | 6.58 | Leaderboard row quoted by Fermion |
| [Nemotron 3.5 ASR Streaming 0.6B](/models/nemotron-3-5-asr-streaming-0-6b) | 2,368 MB | 2.83 | 6.79 | 13.43 | 15.30 | 9.86 | 3.27 | 4.24 | 7.96 | Leaderboard row quoted by Fermion |

## Phonon-2 throughput by hardware

Fermion-reported throughput in multiples of realtime. Error checks use different test samples across hardware; the final column gives fast-path WER against the exact path where both are reported.

| Hardware | Throughput | Run condition | Word error check |
| --- | --- | --- | --- |
| Apple M5 MacBook Air, GPU (MLX) | 174x | Single stream | 2.94% vs 2.94%; 400 LibriSpeech utterances |
| Apple M5 MacBook Air, CPU | 40x | Single stream | 2.33%; 40-clip check |
| Linux x86-64, eight Zen 5 cores (16 vCPU) | 142.8x | Single stream | 3.94% vs 3.91%; full LibriSpeech test-other |
| Linux Arm, eight Google Axion cores | 52.2x | Single stream | 3.90% vs 3.91%; full LibriSpeech test-other |
| Windows x64, 8 vCPU | 21.0x | Single stream | 2.21%; 40-clip check |
| NVIDIA A100 80 GB | 267x / 3,614x | Single stream / batch 128 | 5.20-5.22% vs 5.20%; seven sets |
| NVIDIA H100 80 GB | 465x / 6,680x | Single stream / batch 128 | 5.20-5.22% vs 5.20%; seven sets |

## Runtimes on the same M5 MacBook Air

Fermion-reported throughput on the same 20 dictations, 797 seconds of speech, one stream at a time, model load excluded, each runtime at its defaults.

| Runtime | Throughput |
| --- | --- |
| [Phonon-2 (MLX)](/models/phonon-2) | 174.0x |
| [FluidAudio, Parakeet TDT 0.6B v3 (Core ML)](/models/parakeet-tdt-0-6b-v3) | 104.9x |
| [FluidAudio, Parakeet Redux (Core ML)](/models/parakeet-redux) | 27.8x |
| [Moonshine Tiny](/models/moonshine-tiny) | 26.7x |
| [whisper.cpp, large-v3-turbo (Metal)](/models/whisper-large-v3-turbo) | 17.0x |
| [sherpa-onnx, Parakeet TDT 0.6B v3 (int8)](/models/parakeet-tdt-0-6b-v3) | 16.5x |

## Interpretation limit

Fermion ran Phonon-2, Phonon-1, and Parakeet Redux with the Open ASR Leaderboard code on full test sets. The other accuracy rows are leaderboard figures quoted by Fermion, without independent re-verification here. Some model cards report different runs; those values are not mixed into this comparison. Voxtral's download size is an estimate at 16 bits. Throughput includes log-mel processing and excludes model load; batch-128 GPU results are separate from single-stream results. The Mac runtime comparison uses 20 dictations totaling 797 seconds at each runtime's defaults. Download size does not measure runtime memory. These results do not enter text-model or voice-agent rankings.

## Primary sources

- [Code](https://github.com/fermionresearch/phonon)
- [Owner page](https://www.fermionresearch.com/research/phonon-2/)
- [Phonon-2 model card](https://huggingface.co/FermionResearch/Phonon-2)
- [Fermion models](https://huggingface.co/FermionResearch)

Canonical page: https://benchlm.ai/voice-benchmarks/phonon-2-launch-evaluation
