Official provider results
Qwen3-ASR model evaluation
Qwen’s model card reports word error rates for Qwen3-ASR 1.7B and 0.6B across AMI, Earnings22, GigaSpeech, LibriSpeech, SPGISpeech, and VoxPopuli.
Source snapshots refreshed
Five measurement lanesVOICE / S2S
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Benchmark profile
- Scope
- Mean WER over seven English test sets plus per-set values
- Primary metric
- Word error rate (lower is better)
- Owner
- Qwen
- Available evidence
- Exact figures published on the official model card
What it measures
Voice experience
Is the exchange natural, robust, and responsive?
Available results
- Mean WER (1.7B)
- 5.59
- Mean WER (0.6B)
- 6.31
Provider-run
AMI 9.26, Earnings22 9.88, GigaSpeech 7.25, LS Clean 1.24, LS Other 2.92, SPGISpeech 2.58, VoxPopuli 5.99.
Provider-run
AMI 10.57, Earnings22 10.72, GigaSpeech 7.65, LS Clean 1.69, LS Other 3.97, SPGISpeech 2.74, VoxPopuli 6.80.
Interpretation limit
Provider-run evaluation on public English test sets.