# Audio Realism Benchmark: Voice Benchmark Profile

> Blind listening comparisons measure how human text-to-speech output sounds across phone-agent, conversational, and explainer speech.

- Scope: 500 held-out American-English prompts; female and male voices; human recordings included in the comparison pool
- Measurement lanes: Voice experience
- Primary metric: Bradley–Terry Elo-style rating, win rate, and uncertainty
- Available evidence: Machine-readable public leaderboard
- Owner: Design Arena
- Source snapshots refreshed: 2026-09-18

## Interpretation limit

The leaderboard measures perceived human-likeness for the selected American-English voices and prompt mix. It does not measure factual accuracy, task completion, full-duplex interaction, or multilingual quality.

## Primary sources

- [Owner page](https://www.designarena.ai/methodology/audio-realism-benchmark)

Canonical page: https://benchlm.ai/voice-benchmarks/audio-realism-benchmark
