Skip to main content
BenchLM
Official provider results

Phonon-2 launch evaluation

Phonon-2 has a reported 5.21% mean word error rate across seven English test sets. Fermion's launch report compares eight speech models and measures throughput across hardware and Mac runtimes.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
LibriSpeech test-clean and test-other, AMI, Earnings-22, GigaSpeech, SPGISpeech, and VoxPopuli; Phonon-2 hardware throughput; six runtime configurations on one M5 MacBook Air
Primary metric
Word error rate in percent (lower is better); audio seconds per wall-clock second (higher is better)
Owner
Fermion Research
Available evidence
Exact tables in the launch report, retrieved September 30, 2026; its accuracy table is also on the official model card

What it measures

Voice experience
Is the exchange natural, robust, and responsive?
Latency
How long do responses, tool calls, and full tasks take?

Available results

English transcription accuracy

Word error rate (%), lower is better. Published averages are retained as reported. Download sizes describe weight files.

Word error rate (%), lower is better. Published averages are retained as reported. Download sizes describe weight files.
ModelDownloadLS cleanLS otherAMIEarnings-22GigaSpeechSPGISpeechVoxPopuliMean WEREvidence
Phonon-2164 MB1.723.929.376.968.353.702.465.21Fermion-run
Parakeet TDT 0.6B v32,508 MB1.523.139.425.857.993.633.194.96Leaderboard row quoted by Fermion
Parakeet Redux178 MB1.944.359.167.908.624.013.875.69Fermion-run
Phonon-1415 MB2.115.0310.3112.348.733.673.736.56Fermion-run
Canary 180M Flash737 MB1.523.4212.098.338.872.043.575.69Leaderboard row quoted by Fermion
Voxtral Mini 4B Realtime 2602About 8,000 MB (estimated)1.624.9413.349.318.802.232.606.12Leaderboard row quoted by Fermion
Whisper large-v3-turbo1,618 MB2.133.7113.888.098.472.797.026.58Leaderboard row quoted by Fermion
Nemotron 3.5 ASR Streaming 0.6B2,368 MB2.836.7913.4315.309.863.274.247.96Leaderboard row quoted by Fermion

Phonon-2 throughput by hardware

Fermion-reported throughput in multiples of realtime. Error checks use different test samples across hardware; the final column gives fast-path WER against the exact path where both are reported.

Fermion-reported throughput in multiples of realtime. Error checks use different test samples across hardware; the final column gives fast-path WER against the exact path where both are reported.
HardwareThroughputRun conditionWord error check
Apple M5 MacBook Air, GPU (MLX)174xSingle stream2.94% vs 2.94%; 400 LibriSpeech utterances
Apple M5 MacBook Air, CPU40xSingle stream2.33%; 40-clip check
Linux x86-64, eight Zen 5 cores (16 vCPU)142.8xSingle stream3.94% vs 3.91%; full LibriSpeech test-other
Linux Arm, eight Google Axion cores52.2xSingle stream3.90% vs 3.91%; full LibriSpeech test-other
Windows x64, 8 vCPU21.0xSingle stream2.21%; 40-clip check
NVIDIA A100 80 GB267x / 3,614xSingle stream / batch 1285.20-5.22% vs 5.20%; seven sets
NVIDIA H100 80 GB465x / 6,680xSingle stream / batch 1285.20-5.22% vs 5.20%; seven sets

Runtimes on the same M5 MacBook Air

Fermion-reported throughput on the same 20 dictations, 797 seconds of speech, one stream at a time, model load excluded, each runtime at its defaults.

Fermion-reported throughput on the same 20 dictations, 797 seconds of speech, one stream at a time, model load excluded, each runtime at its defaults.
RuntimeThroughput
Phonon-2 (MLX)174.0x
FluidAudio, Parakeet TDT 0.6B v3 (Core ML)104.9x
FluidAudio, Parakeet Redux (Core ML)27.8x
Moonshine Tiny26.7x
whisper.cpp, large-v3-turbo (Metal)17.0x
sherpa-onnx, Parakeet TDT 0.6B v3 (int8)16.5x

Interpretation limit

Fermion ran Phonon-2, Phonon-1, and Parakeet Redux with the Open ASR Leaderboard code on full test sets. The other accuracy rows are leaderboard figures quoted by Fermion, without independent re-verification here. Some model cards report different runs; those values are not mixed into this comparison. Voxtral's download size is an estimate at 16 bits. Throughput includes log-mel processing and excludes model load; batch-128 GPU results are separate from single-stream results. The Mac runtime comparison uses 20 dictations totaling 797 seconds at each runtime's defaults. Download size does not measure runtime memory. These results do not enter text-model or voice-agent rankings.