Audio Realism Benchmark
Blind listening comparisons measure how human text-to-speech output sounds across phone-agent, conversational, and explainer speech.
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Benchmark profile
- Scope
- 500 held-out American-English prompts; female and male voices; human recordings included in the comparison pool
- Primary metric
- Bradley–Terry Elo-style rating, win rate, and uncertainty
- Owner
- Design Arena
- Available evidence
- Machine-readable public leaderboard
What it measures
How the protocol works
- 01
Vetted native speakers hear two clips generated from the same prompt and select the one that sounds more human.
- 02
The benchmark uses 500 private, held-out prompts split across phone-agent, conversational, and explainer speech.
- 03
One female and one male American-English voice are selected for each model, and comparisons use the same gender.
- 04
Scores are fitted with a Bradley–Terry model and presented as an Elo-style rating; real human recordings remain in the comparison pool.
Available results
| Rank | Model | Elo-style rating | BT uncertainty | Win rate | Battles | Avg. generation |
|---|---|---|---|---|---|---|
| 1 | Bland Speech v3 Bland AI | 1375 | — | 80.8% | 458 | 1.48s |
| 2 | Kalpa TTS Beta v0.1 Kalpa Labs | 1331 | — | 72.5% | 367 | 2.43s |
| 3 | MAI-Voice-2 Microsoft | 1228 | — | 65.7% | 461 | 1.39s |
| 4 | ElevenLabs Eleven v3 Conversational ElevenLabs | 1204 | — | 62.0% | 284 | 2.90s |
| 5 | Cartesia Sonic 3.5 Cartesia | 1182 | — | 59.9% | 638 | 1.68s |
| 6 | Grok TTS xAI | 1163 | — | 57.2% | 460 | 2.06s |
| 7 | Gemini 2.5 Pro TTS Preview Google | 1116 | — | 51.4% | 479 | 6.40s |
| 8 | MiniMax Speech-02 HD MiniMax | 1107 | — | 49.6% | 468 | 3.04s |
| 9 | GPT Realtime 2 OpenAI | 1089 | — | 43.4% | 175 | 2.50s |
| 10 | Gemini 3.1 Flash TTS Preview Google | 1046 | — | 39.0% | 475 | 4.51s |
| 11 | GPT-4o mini TTS OpenAI | 1019 | — | 38.9% | 453 | 1.85s |
| 12 | Gemini 2.5 Flash TTS Preview Google | 1008 | — | 36.1% | 473 | 4.61s |
| 13 | Lightning v3.1 Pro Smallest AI | 971 | — | 32.0% | 463 | 1.73s |
| 14 | ElevenLabs Eleven v3 ElevenLabs | 964 | — | 29.2% | 463 | 3.04s |
| 15 | Inworld TTS-1.5 Max Inworld | 792 | — | 14.2% | 464 | 4.51s |