Which TTS model sounds most human?
Start with Audio Realism. Its blind pairwise listening protocol covers selected American-English voices across phone-agent, conversational, and explainer prompts.
There is no single defensible voice-agent score yet. These tables keep spoken reasoning, workflow success, conversation dynamics, experience, and latency in their original protocols so the numbers remain interpretable.
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Does the system understand and answer spoken requests?
Can it follow policy, use tools, and finish a workflow?
Can it handle turns, interruptions, ambiguity, and state?
Is the exchange natural, robust, and responsive?
How long do responses, tool calls, and full tasks take?
The tabs are separate leaderboards, not ingredients in a BenchLM ranking. Audio Realism uses blind listening preferences; VoiceBench uses its own overall score; AudioAgentBench shows checks passed; Full-Duplex-Bench reports strict pass@1 and seconds.
| Rank | Model | Elo-style rating | BT uncertainty | Win rate | Battles | Avg. generation |
|---|---|---|---|---|---|---|
| 1 | Bland Speech v3 Bland AI | 1365 | ±23.9 | 82.5% | 372 | 1.68s |
| 2 | MAI-Voice-2 Microsoft | 1224 | ±19.8 | 68.9% | 376 | 1.50s |
| 3 | Grok TTS xAI | 1156 | ±18.8 | 60.9% | 378 | 2.27s |
| 4 | Gemini 2.5 Pro TTS Preview Google | 1109 | ±18.2 | 54.9% | 390 | 6.95s |
| 5 | Cartesia Sonic 3.5 Cartesia | 1096 | ±18.2 | 51.6% | 395 | 1.55s |
| 6 | MiniMax Speech-02 HD MiniMax | 1093 | ±18.6 | 52.4% | 380 | 3.36s |
| 7 | Gemini 3.1 Flash TTS Preview Google | 1037 | ±18.4 | 41.5% | 388 | 4.99s |
| 8 | GPT-4o mini TTS OpenAI | 1011 | ±18.9 | 42.8% | 362 | 2.01s |
| 9 | Gemini 2.5 Flash TTS Preview Google | 985 | ±18.9 | 37.9% | 385 | 5.16s |
| 10 | Lightning v3.1 Pro Smallest AI | 952 | ±19.4 | 34.0% | 380 | 1.94s |
| 11 | ElevenLabs Eleven v3 ElevenLabs | 938 | ±20.1 | 29.1% | 364 | 3.35s |
| 12 | Inworld TTS-1.5 Max Inworld | 763 | ±24.9 | 14.2% | 373 | 4.97s |
Snapshot sources: benchmark-owner leaderboard and run records. Full-Duplex-Bench v3 values are transcribed from the official paper table.
| Benchmark | Spoken reasoning | Task completion | Conversation dynamics | Voice experience | Latency | Evidence |
|---|---|---|---|---|---|---|
| Audio Realism Benchmark Blind listening comparisons measure how human text-to-speech output sounds across phone-agent, conversational, and explainer speech. | Live snapshot Machine-readable public leaderboard | |||||
| VoiceBench A spoken question-answering suite spanning open-ended, knowledge, reasoning, instruction-following, and safety tasks. | Live snapshot Machine-readable public leaderboard | |||||
| AudioAgentBench Evaluates real-time voice agents on policy compliance, tool use, grounding, ambiguity, and state tracking. | Live snapshot Machine-readable public run records | |||||
| Full-Duplex-Bench v3 Tests full-duplex voice agents on multi-tool tasks with human speech, disfluencies, interruptions, and latency measurements. | Paper results Results transcribed from the official paper |
Start with Audio Realism. Its blind pairwise listening protocol covers selected American-English voices across phone-agent, conversational, and explainer prompts.
Start with VoiceBench. Check its architecture and individual task columns before relying on the owner-defined overall score.
Use AudioAgentBench for recent public run aggregates, then compare human and synthesized audio rather than merging them.
Use Full-Duplex-Bench v3. Strict pass@1 and latency answer different questions and should be inspected together.