Phonon-2 launch evaluation
Phonon-2 has a reported 5.21% mean word error rate across seven English test sets. Fermion's launch report compares eight speech models and measures throughput across hardware and Mac runtimes.
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Benchmark profile
- Scope
- LibriSpeech test-clean and test-other, AMI, Earnings-22, GigaSpeech, SPGISpeech, and VoxPopuli; Phonon-2 hardware throughput; six runtime configurations on one M5 MacBook Air
- Primary metric
- Word error rate in percent (lower is better); audio seconds per wall-clock second (higher is better)
- Owner
- Fermion Research
- Available evidence
- Exact tables in the launch report, retrieved September 30, 2026; its accuracy table is also on the official model card
What it measures
Available results
English transcription accuracy
Word error rate (%), lower is better. Published averages are retained as reported. Download sizes describe weight files.
| Model | Download | LS clean | LS other | AMI | Earnings-22 | GigaSpeech | SPGISpeech | VoxPopuli | Mean WER | Evidence |
|---|---|---|---|---|---|---|---|---|---|---|
| Phonon-2 | 164 MB | 1.72 | 3.92 | 9.37 | 6.96 | 8.35 | 3.70 | 2.46 | 5.21 | Fermion-run |
| Parakeet TDT 0.6B v3 | 2,508 MB | 1.52 | 3.13 | 9.42 | 5.85 | 7.99 | 3.63 | 3.19 | 4.96 | Leaderboard row quoted by Fermion |
| Parakeet Redux | 178 MB | 1.94 | 4.35 | 9.16 | 7.90 | 8.62 | 4.01 | 3.87 | 5.69 | Fermion-run |
| Phonon-1 | 415 MB | 2.11 | 5.03 | 10.31 | 12.34 | 8.73 | 3.67 | 3.73 | 6.56 | Fermion-run |
| Canary 180M Flash | 737 MB | 1.52 | 3.42 | 12.09 | 8.33 | 8.87 | 2.04 | 3.57 | 5.69 | Leaderboard row quoted by Fermion |
| Voxtral Mini 4B Realtime 2602 | About 8,000 MB (estimated) | 1.62 | 4.94 | 13.34 | 9.31 | 8.80 | 2.23 | 2.60 | 6.12 | Leaderboard row quoted by Fermion |
| Whisper large-v3-turbo | 1,618 MB | 2.13 | 3.71 | 13.88 | 8.09 | 8.47 | 2.79 | 7.02 | 6.58 | Leaderboard row quoted by Fermion |
| Nemotron 3.5 ASR Streaming 0.6B | 2,368 MB | 2.83 | 6.79 | 13.43 | 15.30 | 9.86 | 3.27 | 4.24 | 7.96 | Leaderboard row quoted by Fermion |
Phonon-2 throughput by hardware
Fermion-reported throughput in multiples of realtime. Error checks use different test samples across hardware; the final column gives fast-path WER against the exact path where both are reported.
| Hardware | Throughput | Run condition | Word error check |
|---|---|---|---|
| Apple M5 MacBook Air, GPU (MLX) | 174x | Single stream | 2.94% vs 2.94%; 400 LibriSpeech utterances |
| Apple M5 MacBook Air, CPU | 40x | Single stream | 2.33%; 40-clip check |
| Linux x86-64, eight Zen 5 cores (16 vCPU) | 142.8x | Single stream | 3.94% vs 3.91%; full LibriSpeech test-other |
| Linux Arm, eight Google Axion cores | 52.2x | Single stream | 3.90% vs 3.91%; full LibriSpeech test-other |
| Windows x64, 8 vCPU | 21.0x | Single stream | 2.21%; 40-clip check |
| NVIDIA A100 80 GB | 267x / 3,614x | Single stream / batch 128 | 5.20-5.22% vs 5.20%; seven sets |
| NVIDIA H100 80 GB | 465x / 6,680x | Single stream / batch 128 | 5.20-5.22% vs 5.20%; seven sets |
Runtimes on the same M5 MacBook Air
Fermion-reported throughput on the same 20 dictations, 797 seconds of speech, one stream at a time, model load excluded, each runtime at its defaults.
| Runtime | Throughput |
|---|---|
| Phonon-2 (MLX) | 174.0x |
| FluidAudio, Parakeet TDT 0.6B v3 (Core ML) | 104.9x |
| FluidAudio, Parakeet Redux (Core ML) | 27.8x |
| Moonshine Tiny | 26.7x |
| whisper.cpp, large-v3-turbo (Metal) | 17.0x |
| sherpa-onnx, Parakeet TDT 0.6B v3 (int8) | 16.5x |