Skip to main content
Live source snapshot

Audio Realism Benchmark

Blind listening comparisons measure how human text-to-speech output sounds across phone-agent, conversational, and explainer speech.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
500 held-out American-English prompts; female and male voices; human recordings included in the comparison pool
Primary metric
Bradley–Terry Elo-style rating, win rate, and uncertainty
Owner
Design Arena
Available evidence
Machine-readable public leaderboard

What it measures

Voice experience
Is the exchange natural, robust, and responsive?

How the protocol works

  1. 01

    Vetted native speakers hear two clips generated from the same prompt and select the one that sounds more human.

  2. 02

    The benchmark uses 500 private, held-out prompts split across phone-agent, conversational, and explainer speech.

  3. 03

    One female and one male American-English voice are selected for each model, and comparisons use the same gender.

  4. 04

    Scores are fitted with a Bradley–Terry model and presented as an Elo-style rating; real human recordings remain in the comparison pool.

Available results

Protocol details
RankModelElo-style ratingBT uncertaintyWin rateBattlesAvg. generation
1Bland Speech v3
Bland AI
1365±23.982.5%3721.68s
2MAI-Voice-2
Microsoft
1224±19.868.9%3761.50s
3Grok TTS
xAI
1156±18.860.9%3782.27s
4Gemini 2.5 Pro TTS Preview
Google
1109±18.254.9%3906.95s
5Cartesia Sonic 3.5
Cartesia
1096±18.251.6%3951.55s
6MiniMax Speech-02 HD
MiniMax
1093±18.652.4%3803.36s
7Gemini 3.1 Flash TTS Preview
Google
1037±18.441.5%3884.99s
8GPT-4o mini TTS
OpenAI
1011±18.942.8%3622.01s
9Gemini 2.5 Flash TTS Preview
Google
985±18.937.9%3855.16s
10Lightning v3.1 Pro
Smallest AI
952±19.434.0%3801.94s
11ElevenLabs Eleven v3
ElevenLabs
938±20.129.1%3643.35s
12Inworld TTS-1.5 Max
Inworld
763±24.914.2%3734.97s

Interpretation limit

The leaderboard measures perceived human-likeness for the selected American-English voices and prompt mix. It does not measure factual accuracy, task completion, full-duplex interaction, or multilingual quality.