Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes
Benchmark directory

A map of the voice benchmark landscape

Each benchmark covers a different slice of voice-system quality. The matrix makes those boundaries visible before you compare any scores.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

41 tracked benchmarks

Voice benchmarks with the capability lanes each one covers and the evidence status.
BenchmarkSpoken reasoningTask completionConversation dynamicsVoice experienceLatencyEvidence
Grok Voice Transcribe 2.0 launch evaluation

xAI’s September 18, 2026 launch post reports a rank on an independent streaming speech-to-text leaderboard, four internal word-error-rate sets drawn from production traffic, and one exact multilingual short-phrase pair against Grok Voice Transcribe 1.0.

Provider results

The short-phrase pair and the pricing are stated in the post’s prose. The four internal charts and the multilingual comparison render client-side, so their per-category values are not published in a readable form.

Qwen3.8-Omni-Flash omni evaluation

Qwen’s September 18, 2026 launch post reports a 32-row audio and audio-visual table for Qwen3.8-Omni-Flash against Qwen3.5-Omni-Plus, Gemini 3.8 Flash, Seed 2.0 Lite, and Muse Spark 1.2, plus a separate static-versus-agent comparison on three video sets.

Provider results

Exact figures tabulated in the official launch post for all five compared systems

Canto launch evaluation

Wispr’s September 17, 2026 launch post reports word error rates for Canto against five competing transcription systems on two private dictation sets and three public English datasets.

Provider results

Exact figures labelled on the charts in the official launch post; the two dictation sets are private to Wispr

Gemini 3.8 Audio model evaluation

Google’s September 15 model evaluation reports separate results for Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on spoken agent tasks and a speech-to-speech index. The same-day launch announcement adds a Big Bench Audio score and a Speech Agent Arena standing.

Provider results

Exact bar-chart labels in Google’s model evaluation PDF, plus figures stated in the prose of the launch announcement; EVA-Bench scatterplot points are not tabulated

GPT Audio 1.5 model documentation

OpenAI’s model page documents the modalities, limits, and pricing of gpt-audio-1.5; it publishes no benchmark scores for the model.

Provider results

Specifications only; OpenAI publishes no evaluation numbers for this model

GPT Transcribe model documentation

OpenAI’s model page documents endpoints, features, and per-minute pricing for gpt-transcribe; it publishes no word-error-rate table on the page.

Provider results

Specifications only

GPT Live Transcribe model documentation

OpenAI’s model page documents the realtime-only endpoint, latency controls, and per-minute pricing for gpt-live-transcribe; it publishes no accuracy table on the page.

Provider results

Specifications only

Gemini 3.5 Transcribe launch evaluation

Google’s August 26, 2026 launch post reports word error rates on an independent transcription leaderboard and FLEURS for pre-recorded and streaming transcription, plus a latency gain over Chirp 3.

Provider results

Exact figures published in the official launch post

MAI-Transcribe-2 launch evaluation

Microsoft AI’s September 3, 2026 launch post reports FLEURS word error rate across 60 languages, an independent leaderboard position, and relative speed against competing transcription models.

Provider results

Exact figures published in the official launch post

MAI-Transcribe-1.5 launch evaluation

Microsoft’s June 2, 2026 Foundry post reports the FLEURS word error rate improvement for MAI-Transcribe-1.5 over MAI-Transcribe-1.

Provider results

Exact figures published in the official Foundry post

MAI-Voice-2-Flash launch notes

Microsoft AI’s July 23, 2026 post positions MAI-Voice-2-Flash as a faster, cheaper sibling of MAI-Voice-2 without publishing benchmark scores.

Provider results

Relative claims only

Sonic 3.6 launch evaluation

Cartesia’s August 27, 2026 launch post reports independent controlled-voice Elo ratings, blind preference win rates against Eleven v3, response latency, and locale coverage for Sonic 3.6.

Provider results

Exact figures published in the official launch post

Realtime TTS-2 launch evaluation

Inworld’s launch post reports an independent speech-arena ranking, time-to-first-audio, and language coverage for Realtime TTS-2.

Provider results

Rank and latency published in the official launch post and docs; no Elo value given

Realtime TTS-2 Flash launch notes

Inworld’s model documentation reports time-to-first-byte and coverage for the Flash variant of Realtime TTS-2.

Provider results

Documented figures only

SeedRealtime launch notes

ByteDance Seed’s August 5, 2026 announcement describes an end-to-end audio-visual full-duplex model and reports a halving of conversational pacing issues versus cascaded pipelines.

Provider results

Relative claim only

Hy ASR 3.0 preview launch notes

Tencent’s Hunyuan site describes Hy ASR 3.0 preview as a Hy3-based recogniser with context-aware correction and broad dialect coverage; launch coverage cites word error rates near 3%.

Provider results

Tencent’s site is descriptive; the WER values come from launch coverage and are marked secondary

Deepgram Flux launch evaluation

Deepgram’s October 2, 2025 launch announcement reports end-of-turn detection latency and Nova-3-level accuracy for Flux.

Provider results

Latency figure published; accuracy stated relative to Nova-3 without a WER value

Flux Multilingual launch evaluation

Deepgram’s April 29, 2026 press release reports end-of-turn latency and language coverage for Flux Multilingual.

Provider results

Latency and coverage published; accuracy stated as monolingual-grade without a WER value

Pocket TTS project notes

Kyutai’s TTS page and repository document Pocket TTS’s size, CPU real-time factor, and first-audio latency.

Provider results

Documented figures only; no listening-test result

Zonos 2 model evaluation

Zyphra’s ZONOS2 GGUF model card reports word error rate, speaker similarity, and UTMOS for the reference F16 build.

Provider results

Exact figures published on the official GGUF model card

Granite Speech 5.0 TurboCTC model notes

IBM’s model card and Granite 4.2 blog document the model’s size, training data, and throughput; Open ASR leaderboard results are published as images.

Provider results

Throughput published in text; WER charts are images

Cohere Transcribe Arabic model notes

Cohere’s catalog and Hugging Face repository document the Arabic transcription model’s lineage and availability; no evaluation table is published.

Provider results

No evaluation numbers published

VibeVoice-ASR-Streaming model notes

Microsoft’s model card documents streaming speaker-attributed transcription, hotword support, and ten-language coverage; evaluation results are published as an image.

Provider results

Evaluation table published only as an image

Qwen3-ASR model evaluation

Qwen’s model card reports word error rates for Qwen3-ASR 1.7B and 0.6B across AMI, Earnings22, GigaSpeech, LibriSpeech, SPGISpeech, and VoxPopuli.

Provider results

Exact figures published on the official model card

AuK model notes

Tencent’s model card documents AuK’s task coverage across speech generation, editing, enhancement, and separation; its benchmark chart is published as an image.

Provider results

Benchmark chart published only as an image

GPT-Live-1 API launch evaluation

OpenAI’s API launch post reports spoken task success, full-duplex interactivity, tool calling, response quality, and turn-taking latency for GPT-Live-1 against GPT-Realtime-2.1 and GPT-Realtime-2.

Provider results

Exact figures embedded in the official launch post charts

Muse Voice Transcribe launch evaluation

Meta’s launch evaluation reports streaming transcription accuracy, speaker diarization, and real-time audio-perception capabilities for Muse Voice Transcribe.

Provider results

Exact figures published in the official launch post

Artificial Analysis Speech-to-Speech Index

A source-owned index for native audio models, with separate results for spoken reasoning, customer-service task completion, live preference, conversation flow, latency, and audio cost.

Live snapshot

Published source table and repeatable page snapshot

Big Bench Audio

One thousand spoken reasoning questions test whether native audio models can answer correctly from audio and how quickly they begin replying.

Live snapshot

Published source table and public audio dataset

Speech Agent Arena

People hold blind, paired voice conversations and choose the system they prefer; eligible tool-calling conversations also receive a task-success check.

Live snapshot

Published Arena leaderboard with confidence intervals and sample counts

Audio Realism Benchmark

Blind listening comparisons measure how human text-to-speech output sounds across phone-agent, conversational, and explainer speech.

Live snapshot

Machine-readable public leaderboard

VoiceBench

A spoken question-answering suite spanning open-ended, knowledge, reasoning, instruction-following, and safety tasks.

Live snapshot

Machine-readable public leaderboard

MMAU

Measures expert-level understanding and reasoning across speech, environmental sound, and music.

Live snapshot

Machine-readable benchmark-owner leaderboard

MMAU-Pro

Tests harder audio reasoning across speech, sound, music, spatial audio, multiple clips, voice chat, and instruction following.

Live snapshot

Machine-readable benchmark-owner leaderboard

AudioAgentBench

Evaluates real-time voice agents on policy compliance, tool use, grounding, ambiguity, and state tracking.

Live snapshot

Machine-readable public run records

Full-Duplex-Bench v3

Tests full-duplex voice agents on multi-tool tasks with human speech, disfluencies, interruptions, and latency measurements.

Paper results

Results transcribed from the official paper

τ-Voice

Extends τ-bench into spoken customer-service tasks to measure the capability gap between text and voice agents.

Live snapshot

Machine-readable benchmark-owner submissions and reproducible evaluation code

EVA-Bench

Separates enterprise task accuracy from conversational experience across realistic service scenarios.

Live snapshot

Machine-readable benchmark-owner leaderboard, code, and public dataset

Voice Agent Latency Benchmark

Measures the silence callers experience between finishing a turn and hearing a voice-agent platform begin its reply.

Live snapshot

Machine-readable public API and reproducible raw-call artifacts

VoiceAgentBench

A multilingual voice-agent suite covering tools, workflows, multi-turn interaction, and safety.

Reference

Paper and public dataset

ADU-Bench

Measures how voice systems handle ambiguous, disfluent, and underspecified spoken dialogue.

Reference

Paper and benchmark resources

Measurement lanes

01

Spoken reasoning

Does the system understand and answer spoken requests?

02

Task completion

Can it follow policy, use tools, and finish a workflow?

03

Conversation dynamics

Can it handle turns, interruptions, ambiguity, and state?

04

Voice experience

Is the exchange natural, robust, and responsive?

05

Latency

How long do responses, tool calls, and full tasks take?