Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes
Voice systems

Voice benchmarks, separated by what they measure

There is no single defensible voice-agent score yet. These tables keep spoken reasoning, workflow success, conversation dynamics, experience, and latency in their original protocols so the numbers remain interpretable.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

03

Spoken reasoning

Does the system understand and answer spoken requests?

06

Task completion

Can it follow policy, use tools, and finish a workflow?

13

Conversation dynamics

Can it handle turns, interruptions, ambiguity, and state?

17

Voice experience

Is the exchange natural, robust, and responsive?

17

Latency

How long do responses, tool calls, and full tasks take?

Current source-backed results

The tabs are separate leaderboards, not ingredients in a BenchLM ranking. Audio Realism uses blind listening preferences; VoiceBench uses its own overall score; AudioAgentBench shows checks passed; Full-Duplex-Bench reports strict pass@1 and seconds.

Protocol details
Audio-realism arena results: Elo-style rating, uncertainty, win rate, battles, and generation time per model.
RankModelElo-style ratingBT uncertaintyWin rateBattlesAvg. generation
1Bland Speech v3
Bland AI
137580.8%4581.48s
2Kalpa TTS Beta v0.1
Kalpa Labs
133172.5%3672.43s
3MAI-Voice-2
Microsoft
122865.7%4611.39s
4ElevenLabs Eleven v3 Conversational
ElevenLabs
120462.0%2842.90s
5Cartesia Sonic 3.5
Cartesia
118259.9%6381.68s
6Grok TTS
xAI
116357.2%4602.06s
7Gemini 2.5 Pro TTS Preview
Google
111651.4%4796.40s
8MiniMax Speech-02 HD
MiniMax
110749.6%4683.04s
9GPT Realtime 2
OpenAI
108943.4%1752.50s
10Gemini 3.1 Flash TTS Preview
Google
104639.0%4754.51s
11GPT-4o mini TTS
OpenAI
101938.9%4531.85s
12Gemini 2.5 Flash TTS Preview
Google
100836.1%4734.61s
13Lightning v3.1 Pro
Smallest AI
97132.0%4631.73s
14ElevenLabs Eleven v3
ElevenLabs
96429.2%4633.04s
15Inworld TTS-1.5 Max
Inworld
79214.2%4644.51s

Snapshot sources: benchmark-owner leaderboard and run records. Full-Duplex-Bench v3 values are transcribed from the official paper table.

Benchmark coverage

Voice benchmarks with the capability lanes each one covers and the evidence status.
BenchmarkSpoken reasoningTask completionConversation dynamicsVoice experienceLatencyEvidence
GPT Audio 1.5 model documentation

OpenAI’s model page documents the modalities, limits, and pricing of gpt-audio-1.5; it publishes no benchmark scores for the model.

Provider results

Specifications only; OpenAI publishes no evaluation numbers for this model

GPT Transcribe model documentation

OpenAI’s model page documents endpoints, features, and per-minute pricing for gpt-transcribe; it publishes no word-error-rate table on the page.

Provider results

Specifications only

GPT Live Transcribe model documentation

OpenAI’s model page documents the realtime-only endpoint, latency controls, and per-minute pricing for gpt-live-transcribe; it publishes no accuracy table on the page.

Provider results

Specifications only

Gemini 3.5 Transcribe launch evaluation

Google’s August 26, 2026 launch post reports word error rates on an independent transcription leaderboard and FLEURS for pre-recorded and streaming transcription, plus a latency gain over Chirp 3.

Provider results

Exact figures published in the official launch post

What the current results can answer

Which TTS model sounds most human?

Start with Audio Realism. Its blind pairwise listening protocol covers selected American-English voices across phone-agent, conversational, and explainer prompts.

Which model answers spoken questions best?

Start with VoiceBench. Check its architecture and individual task columns before relying on the owner-defined overall score.

Which agent completes service workflows?

Use AudioAgentBench for recent public run aggregates, then compare human and synthesized audio rather than merging them.

Which full-duplex system balances success and speed?

Use Full-Duplex-Bench v3. Strict pass@1 and latency answer different questions and should be inspected together.