Official provider results
Zonos 2 model evaluation
Zyphra’s ZONOS2 GGUF model card reports word error rate, speaker similarity, and UTMOS for the reference F16 build.
Source snapshots refreshed
Five measurement lanesVOICE / S2S
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Benchmark profile
- Scope
- WER, speaker similarity, and UTMOS naturalness for the F16 reference build on Zyphra’s evaluation set
- Primary metric
- WER (lower is better), speaker similarity and UTMOS (higher is better)
- Owner
- Zyphra
- Available evidence
- Exact figures published on the official GGUF model card
What it measures
Voice experience
Is the exchange natural, robust, and responsive?
Available results
- WER (F16 reference)
- 2.79
- Speaker similarity
- 66.75
- UTMOS
- 4.40
Provider-run
Q8_0 build 2.87; lower is better.
Provider-run
Voice-cloning fidelity; higher is better.
Provider-run
Predicted naturalness; higher is better.
Interpretation limit
Provider-run evaluation on Zyphra’s own set; the blog compares against other systems qualitatively only.