Which TTS model sounds most human?
Start with Audio Realism. Its blind pairwise listening protocol covers selected American-English voices across phone-agent, conversational, and explainer prompts.
Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.
Follow model changesNo single score covers every voice-agent workload. The source-owned speech-to-speech index combines four native-audio checks; the other tables keep reasoning, workflow success, conversation dynamics, experience, and latency in their own protocols.
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Does the system understand and answer spoken requests?
Can it follow policy, use tools, and finish a workflow?
Can it handle turns, interruptions, ambiguity, and state?
Is the exchange natural, robust, and responsive?
How long do responses, tool calls, and full tasks take?
The tabs are separate leaderboards, not ingredients in a BenchLM ranking. Audio Realism uses blind listening preferences; VoiceBench uses its own overall score; AudioAgentBench shows checks passed; Full-Duplex-Bench reports strict pass@1 and seconds.
| Rank | Model | Elo-style rating | BT uncertainty | Win rate | Battles | Avg. generation |
|---|---|---|---|---|---|---|
| 1 | Bland Speech v3 Bland AI | 1375 | — | 80.8% | 458 | 1.48s |
| 2 | Kalpa TTS Beta v0.1 Kalpa Labs | 1331 | — | 72.5% | 367 | 2.43s |
| 3 | MAI-Voice-2 Microsoft | 1228 | — | 65.7% | 461 | 1.39s |
| 4 | ElevenLabs Eleven v3 Conversational ElevenLabs | 1204 | — | 62.0% | 284 | 2.90s |
| 5 | Cartesia Sonic 3.5 Cartesia | 1182 | — | 59.9% | 638 | 1.68s |
| 6 | Grok TTS xAI | 1163 | — | 57.2% | 460 | 2.06s |
| 7 | Gemini 2.5 Pro TTS Preview Google | 1116 | — | 51.4% | 479 | 6.40s |
| 8 | MiniMax Speech-02 HD MiniMax | 1107 | — | 49.6% | 468 | 3.04s |
| 9 | GPT Realtime 2 OpenAI | 1089 | — | 43.4% | 175 | 2.50s |
| 10 | Gemini 3.1 Flash TTS Preview Google | 1046 | — | 39.0% | 475 | 4.51s |
| 11 | GPT-4o mini TTS OpenAI | 1019 | — | 38.9% | 453 | 1.85s |
| 12 | Gemini 2.5 Flash TTS Preview Google | 1008 | — | 36.1% | 473 | 4.61s |
| 13 | Lightning v3.1 Pro Smallest AI | 971 | — | 32.0% | 463 | 1.73s |
| 14 | ElevenLabs Eleven v3 ElevenLabs | 964 | — | 29.2% | 463 | 3.04s |
| 15 | Inworld TTS-1.5 Max Inworld | 792 | — | 14.2% | 464 | 4.51s |
The source index covers native audio models with reasoning, τ-Voice task success, Arena preference, and eligible tool-call success. Its component results remain visible in separate profiles.
Snapshot sources: benchmark-owner leaderboard and run records. Full-Duplex-Bench v3 values are transcribed from the official paper table.
| Benchmark | Spoken reasoning | Task completion | Conversation dynamics | Voice experience | Latency | Evidence |
|---|---|---|---|---|---|---|
| Grok Voice Transcribe 2.0 launch evaluation xAI’s September 18, 2026 launch post reports a rank on an independent streaming speech-to-text leaderboard, four internal word-error-rate sets drawn from production traffic, and one exact multilingual short-phrase pair against Grok Voice Transcribe 1.0. | Provider results The short-phrase pair and the pricing are stated in the post’s prose. The four internal charts and the multilingual comparison render client-side, so their per-category values are not published in a readable form. | |||||
| Qwen3.8-Omni-Flash omni evaluation Qwen’s September 18, 2026 launch post reports a 32-row audio and audio-visual table for Qwen3.8-Omni-Flash against Qwen3.5-Omni-Plus, Gemini 3.8 Flash, Seed 2.0 Lite, and Muse Spark 1.2, plus a separate static-versus-agent comparison on three video sets. | Provider results Exact figures tabulated in the official launch post for all five compared systems | |||||
| Canto launch evaluation Wispr’s September 17, 2026 launch post reports word error rates for Canto against five competing transcription systems on two private dictation sets and three public English datasets. | Provider results Exact figures labelled on the charts in the official launch post; the two dictation sets are private to Wispr | |||||
| Gemini 3.8 Audio model evaluation Google’s September 15 model evaluation reports separate results for Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on spoken agent tasks and a speech-to-speech index. The same-day launch announcement adds a Big Bench Audio score and a Speech Agent Arena standing. | Provider results Exact bar-chart labels in Google’s model evaluation PDF, plus figures stated in the prose of the launch announcement; EVA-Bench scatterplot points are not tabulated |
Start with Audio Realism. Its blind pairwise listening protocol covers selected American-English voices across phone-agent, conversational, and explainer prompts.
Start with VoiceBench. Check its architecture and individual task columns before relying on the owner-defined overall score.
Use AudioAgentBench for recent public run aggregates, then compare human and synthesized audio rather than merging them.
Use Full-Duplex-Bench v3. Strict pass@1 and latency answer different questions and should be inspected together.