Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes
Voice systems

Voice benchmarks, separated by what they measure

No single score covers every voice-agent workload. The source-owned speech-to-speech index combines four native-audio checks; the other tables keep reasoning, workflow success, conversation dynamics, experience, and latency in their own protocols.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

06

Spoken reasoning

Does the system understand and answer spoken requests?

10

Task completion

Can it follow policy, use tools, and finish a workflow?

17

Conversation dynamics

Can it handle turns, interruptions, ambiguity, and state?

23

Voice experience

Is the exchange natural, robust, and responsive?

19

Latency

How long do responses, tool calls, and full tasks take?

Current source-backed results

The tabs are separate leaderboards, not ingredients in a BenchLM ranking. Audio Realism uses blind listening preferences; VoiceBench uses its own overall score; AudioAgentBench shows checks passed; Full-Duplex-Bench reports strict pass@1 and seconds.

Protocol details
Audio-realism arena results: Elo-style rating, uncertainty, win rate, battles, and generation time per model.
RankModelElo-style ratingBT uncertaintyWin rateBattlesAvg. generation
1Bland Speech v3
Bland AI
137580.8%4581.48s
2Kalpa TTS Beta v0.1
Kalpa Labs
133172.5%3672.43s
3MAI-Voice-2
Microsoft
122865.7%4611.39s
4ElevenLabs Eleven v3 Conversational
ElevenLabs
120462.0%2842.90s
5Cartesia Sonic 3.5
Cartesia
118259.9%6381.68s
6Grok TTS
xAI
116357.2%4602.06s
7Gemini 2.5 Pro TTS Preview
Google
111651.4%4796.40s
8MiniMax Speech-02 HD
MiniMax
110749.6%4683.04s
9GPT Realtime 2
OpenAI
108943.4%1752.50s
10Gemini 3.1 Flash TTS Preview
Google
104639.0%4754.51s
11GPT-4o mini TTS
OpenAI
101938.9%4531.85s
12Gemini 2.5 Flash TTS Preview
Google
100836.1%4734.61s
13Lightning v3.1 Pro
Smallest AI
97132.0%4631.73s
14ElevenLabs Eleven v3
ElevenLabs
96429.2%4633.04s
15Inworld TTS-1.5 Max
Inworld
79214.2%4644.51s

Artificial Analysis speech-to-speech results

The source index covers native audio models with reasoning, τ-Voice task success, Arena preference, and eligible tool-call success. Its component results remain visible in separate profiles.

Snapshot sources: benchmark-owner leaderboard and run records. Full-Duplex-Bench v3 values are transcribed from the official paper table.

Benchmark coverage

Voice benchmarks with the capability lanes each one covers and the evidence status.
BenchmarkSpoken reasoningTask completionConversation dynamicsVoice experienceLatencyEvidence
Grok Voice Transcribe 2.0 launch evaluation

xAI’s September 18, 2026 launch post reports a rank on an independent streaming speech-to-text leaderboard, four internal word-error-rate sets drawn from production traffic, and one exact multilingual short-phrase pair against Grok Voice Transcribe 1.0.

Provider results

The short-phrase pair and the pricing are stated in the post’s prose. The four internal charts and the multilingual comparison render client-side, so their per-category values are not published in a readable form.

Qwen3.8-Omni-Flash omni evaluation

Qwen’s September 18, 2026 launch post reports a 32-row audio and audio-visual table for Qwen3.8-Omni-Flash against Qwen3.5-Omni-Plus, Gemini 3.8 Flash, Seed 2.0 Lite, and Muse Spark 1.2, plus a separate static-versus-agent comparison on three video sets.

Provider results

Exact figures tabulated in the official launch post for all five compared systems

Canto launch evaluation

Wispr’s September 17, 2026 launch post reports word error rates for Canto against five competing transcription systems on two private dictation sets and three public English datasets.

Provider results

Exact figures labelled on the charts in the official launch post; the two dictation sets are private to Wispr

Gemini 3.8 Audio model evaluation

Google’s September 15 model evaluation reports separate results for Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on spoken agent tasks and a speech-to-speech index. The same-day launch announcement adds a Big Bench Audio score and a Speech Agent Arena standing.

Provider results

Exact bar-chart labels in Google’s model evaluation PDF, plus figures stated in the prose of the launch announcement; EVA-Bench scatterplot points are not tabulated

What the current results can answer

Which TTS model sounds most human?

Start with Audio Realism. Its blind pairwise listening protocol covers selected American-English voices across phone-agent, conversational, and explainer prompts.

Which model answers spoken questions best?

Start with VoiceBench. Check its architecture and individual task columns before relying on the owner-defined overall score.

Which agent completes service workflows?

Use AudioAgentBench for recent public run aggregates, then compare human and synthesized audio rather than merging them.

Which full-duplex system balances success and speed?

Use Full-Duplex-Bench v3. Strict pass@1 and latency answer different questions and should be inspected together.