Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes
Live source snapshot

Audio Realism Benchmark

Blind listening comparisons measure how human text-to-speech output sounds across phone-agent, conversational, and explainer speech.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
500 held-out American-English prompts; female and male voices; human recordings included in the comparison pool
Primary metric
Bradley–Terry Elo-style rating, win rate, and uncertainty
Owner
Design Arena
Available evidence
Machine-readable public leaderboard

What it measures

Voice experience
Is the exchange natural, robust, and responsive?

How the protocol works

  1. 01

    Vetted native speakers hear two clips generated from the same prompt and select the one that sounds more human.

  2. 02

    The benchmark uses 500 private, held-out prompts split across phone-agent, conversational, and explainer speech.

  3. 03

    One female and one male American-English voice are selected for each model, and comparisons use the same gender.

  4. 04

    Scores are fitted with a Bradley–Terry model and presented as an Elo-style rating; real human recordings remain in the comparison pool.

Available results

Protocol details
Audio-realism arena results: Elo-style rating, uncertainty, win rate, battles, and generation time per model.
RankModelElo-style ratingBT uncertaintyWin rateBattlesAvg. generation
1Bland Speech v3
Bland AI
137580.8%4581.48s
2Kalpa TTS Beta v0.1
Kalpa Labs
133172.5%3672.43s
3MAI-Voice-2
Microsoft
122865.7%4611.39s
4ElevenLabs Eleven v3 Conversational
ElevenLabs
120462.0%2842.90s
5Cartesia Sonic 3.5
Cartesia
118259.9%6381.68s
6Grok TTS
xAI
116357.2%4602.06s
7Gemini 2.5 Pro TTS Preview
Google
111651.4%4796.40s
8MiniMax Speech-02 HD
MiniMax
110749.6%4683.04s
9GPT Realtime 2
OpenAI
108943.4%1752.50s
10Gemini 3.1 Flash TTS Preview
Google
104639.0%4754.51s
11GPT-4o mini TTS
OpenAI
101938.9%4531.85s
12Gemini 2.5 Flash TTS Preview
Google
100836.1%4734.61s
13Lightning v3.1 Pro
Smallest AI
97132.0%4631.73s
14ElevenLabs Eleven v3
ElevenLabs
96429.2%4633.04s
15Inworld TTS-1.5 Max
Inworld
79214.2%4644.51s

Interpretation limit

The leaderboard measures perceived human-likeness for the selected American-English voices and prompt mix. It does not measure factual accuracy, task completion, full-duplex interaction, or multilingual quality.