Which TTS model sounds most human?
Start with Audio Realism. Its blind pairwise listening protocol covers selected American-English voices across phone-agent, conversational, and explainer prompts.
Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.
Follow model changesThere is no single defensible voice-agent score yet. These tables keep spoken reasoning, workflow success, conversation dynamics, experience, and latency in their original protocols so the numbers remain interpretable.
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Does the system understand and answer spoken requests?
Can it follow policy, use tools, and finish a workflow?
Can it handle turns, interruptions, ambiguity, and state?
Is the exchange natural, robust, and responsive?
How long do responses, tool calls, and full tasks take?
The tabs are separate leaderboards, not ingredients in a BenchLM ranking. Audio Realism uses blind listening preferences; VoiceBench uses its own overall score; AudioAgentBench shows checks passed; Full-Duplex-Bench reports strict pass@1 and seconds.
| Rank | Model | Elo-style rating | BT uncertainty | Win rate | Battles | Avg. generation |
|---|---|---|---|---|---|---|
| 1 | Bland Speech v3 Bland AI | 1375 | — | 80.8% | 458 | 1.48s |
| 2 | Kalpa TTS Beta v0.1 Kalpa Labs | 1331 | — | 72.5% | 367 | 2.43s |
| 3 | MAI-Voice-2 Microsoft | 1228 | — | 65.7% | 461 | 1.39s |
| 4 | ElevenLabs Eleven v3 Conversational ElevenLabs | 1204 | — | 62.0% | 284 | 2.90s |
| 5 | Cartesia Sonic 3.5 Cartesia | 1182 | — | 59.9% | 638 | 1.68s |
| 6 | Grok TTS xAI | 1163 | — | 57.2% | 460 | 2.06s |
| 7 | Gemini 2.5 Pro TTS Preview Google | 1116 | — | 51.4% | 479 | 6.40s |
| 8 | MiniMax Speech-02 HD MiniMax | 1107 | — | 49.6% | 468 | 3.04s |
| 9 | GPT Realtime 2 OpenAI | 1089 | — | 43.4% | 175 | 2.50s |
| 10 | Gemini 3.1 Flash TTS Preview Google | 1046 | — | 39.0% | 475 | 4.51s |
| 11 | GPT-4o mini TTS OpenAI | 1019 | — | 38.9% | 453 | 1.85s |
| 12 | Gemini 2.5 Flash TTS Preview Google | 1008 | — | 36.1% | 473 | 4.61s |
| 13 | Lightning v3.1 Pro Smallest AI | 971 | — | 32.0% | 463 | 1.73s |
| 14 | ElevenLabs Eleven v3 ElevenLabs | 964 | — | 29.2% | 463 | 3.04s |
| 15 | Inworld TTS-1.5 Max Inworld | 792 | — | 14.2% | 464 | 4.51s |
Snapshot sources: benchmark-owner leaderboard and run records. Full-Duplex-Bench v3 values are transcribed from the official paper table.
| Benchmark | Spoken reasoning | Task completion | Conversation dynamics | Voice experience | Latency | Evidence |
|---|---|---|---|---|---|---|
| GPT Audio 1.5 model documentation OpenAI’s model page documents the modalities, limits, and pricing of gpt-audio-1.5; it publishes no benchmark scores for the model. | Provider results Specifications only; OpenAI publishes no evaluation numbers for this model | |||||
| GPT Transcribe model documentation OpenAI’s model page documents endpoints, features, and per-minute pricing for gpt-transcribe; it publishes no word-error-rate table on the page. | Provider results Specifications only | |||||
| GPT Live Transcribe model documentation OpenAI’s model page documents the realtime-only endpoint, latency controls, and per-minute pricing for gpt-live-transcribe; it publishes no accuracy table on the page. | Provider results Specifications only | |||||
| Gemini 3.5 Transcribe launch evaluation Google’s August 26, 2026 launch post reports word error rates on an independent transcription leaderboard and FLEURS for pre-recorded and streaming transcription, plus a latency gain over Chirp 3. | Provider results Exact figures published in the official launch post |
Start with Audio Realism. Its blind pairwise listening protocol covers selected American-English voices across phone-agent, conversational, and explainer prompts.
Start with VoiceBench. Check its architecture and individual task columns before relying on the owner-defined overall score.
Use AudioAgentBench for recent public run aggregates, then compare human and synthesized audio rather than merging them.
Use Full-Duplex-Bench v3. Strict pass@1 and latency answer different questions and should be inspected together.