Skip to main content
Voice systems

Voice benchmarks, separated by what they measure

There is no single defensible voice-agent score yet. These tables keep spoken reasoning, workflow success, conversation dynamics, experience, and latency in their original protocols so the numbers remain interpretable.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

01

Spoken reasoning

Does the system understand and answer spoken requests?

05

Task completion

Can it follow policy, use tools, and finish a workflow?

06

Conversation dynamics

Can it handle turns, interruptions, ambiguity, and state?

02

Voice experience

Is the exchange natural, robust, and responsive?

01

Latency

How long do responses, tool calls, and full tasks take?

Current source-backed results

The tabs are separate leaderboards, not ingredients in a BenchLM ranking. Audio Realism uses blind listening preferences; VoiceBench uses its own overall score; AudioAgentBench shows checks passed; Full-Duplex-Bench reports strict pass@1 and seconds.

Protocol details
RankModelElo-style ratingBT uncertaintyWin rateBattlesAvg. generation
1Bland Speech v3
Bland AI
1365±23.982.5%3721.68s
2MAI-Voice-2
Microsoft
1224±19.868.9%3761.50s
3Grok TTS
xAI
1156±18.860.9%3782.27s
4Gemini 2.5 Pro TTS Preview
Google
1109±18.254.9%3906.95s
5Cartesia Sonic 3.5
Cartesia
1096±18.251.6%3951.55s
6MiniMax Speech-02 HD
MiniMax
1093±18.652.4%3803.36s
7Gemini 3.1 Flash TTS Preview
Google
1037±18.441.5%3884.99s
8GPT-4o mini TTS
OpenAI
1011±18.942.8%3622.01s
9Gemini 2.5 Flash TTS Preview
Google
985±18.937.9%3855.16s
10Lightning v3.1 Pro
Smallest AI
952±19.434.0%3801.94s
11ElevenLabs Eleven v3
ElevenLabs
938±20.129.1%3643.35s
12Inworld TTS-1.5 Max
Inworld
763±24.914.2%3734.97s

Snapshot sources: benchmark-owner leaderboard and run records. Full-Duplex-Bench v3 values are transcribed from the official paper table.

Benchmark coverage

BenchmarkSpoken reasoningTask completionConversation dynamicsVoice experienceLatencyEvidence
Audio Realism Benchmark

Blind listening comparisons measure how human text-to-speech output sounds across phone-agent, conversational, and explainer speech.

Live snapshot

Machine-readable public leaderboard

VoiceBench

A spoken question-answering suite spanning open-ended, knowledge, reasoning, instruction-following, and safety tasks.

Live snapshot

Machine-readable public leaderboard

AudioAgentBench

Evaluates real-time voice agents on policy compliance, tool use, grounding, ambiguity, and state tracking.

Live snapshot

Machine-readable public run records

Full-Duplex-Bench v3

Tests full-duplex voice agents on multi-tool tasks with human speech, disfluencies, interruptions, and latency measurements.

Paper results

Results transcribed from the official paper

What the current results can answer

Which TTS model sounds most human?

Start with Audio Realism. Its blind pairwise listening protocol covers selected American-English voices across phone-agent, conversational, and explainer prompts.

Which model answers spoken questions best?

Start with VoiceBench. Check its architecture and individual task columns before relying on the owner-defined overall score.

Which agent completes service workflows?

Use AudioAgentBench for recent public run aggregates, then compare human and synthesized audio rather than merging them.

Which full-duplex system balances success and speed?

Use Full-Duplex-Bench v3. Strict pass@1 and latency answer different questions and should be inspected together.