Spoken reasoning
Does the system understand and answer spoken requests?
Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.
Follow model changesEach benchmark covers a different slice of voice-system quality. The matrix makes those boundaries visible before you compare any scores.
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
| Benchmark | Spoken reasoning | Task completion | Conversation dynamics | Voice experience | Latency | Evidence |
|---|---|---|---|---|---|---|
| Grok Voice Transcribe 2.0 launch evaluation xAI’s September 18, 2026 launch post reports a rank on an independent streaming speech-to-text leaderboard, four internal word-error-rate sets drawn from production traffic, and one exact multilingual short-phrase pair against Grok Voice Transcribe 1.0. | Provider results The short-phrase pair and the pricing are stated in the post’s prose. The four internal charts and the multilingual comparison render client-side, so their per-category values are not published in a readable form. | |||||
| Qwen3.8-Omni-Flash omni evaluation Qwen’s September 18, 2026 launch post reports a 32-row audio and audio-visual table for Qwen3.8-Omni-Flash against Qwen3.5-Omni-Plus, Gemini 3.8 Flash, Seed 2.0 Lite, and Muse Spark 1.2, plus a separate static-versus-agent comparison on three video sets. | Provider results Exact figures tabulated in the official launch post for all five compared systems | |||||
| Canto launch evaluation Wispr’s September 17, 2026 launch post reports word error rates for Canto against five competing transcription systems on two private dictation sets and three public English datasets. | Provider results Exact figures labelled on the charts in the official launch post; the two dictation sets are private to Wispr | |||||
| Gemini 3.8 Audio model evaluation Google’s September 15 model evaluation reports separate results for Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on spoken agent tasks and a speech-to-speech index. The same-day launch announcement adds a Big Bench Audio score and a Speech Agent Arena standing. | Provider results Exact bar-chart labels in Google’s model evaluation PDF, plus figures stated in the prose of the launch announcement; EVA-Bench scatterplot points are not tabulated | |||||
| GPT Audio 1.5 model documentation OpenAI’s model page documents the modalities, limits, and pricing of gpt-audio-1.5; it publishes no benchmark scores for the model. | Provider results Specifications only; OpenAI publishes no evaluation numbers for this model | |||||
| GPT Transcribe model documentation OpenAI’s model page documents endpoints, features, and per-minute pricing for gpt-transcribe; it publishes no word-error-rate table on the page. | Provider results Specifications only | |||||
| GPT Live Transcribe model documentation OpenAI’s model page documents the realtime-only endpoint, latency controls, and per-minute pricing for gpt-live-transcribe; it publishes no accuracy table on the page. | Provider results Specifications only | |||||
| Gemini 3.5 Transcribe launch evaluation Google’s August 26, 2026 launch post reports word error rates on an independent transcription leaderboard and FLEURS for pre-recorded and streaming transcription, plus a latency gain over Chirp 3. | Provider results Exact figures published in the official launch post | |||||
| MAI-Transcribe-2 launch evaluation Microsoft AI’s September 3, 2026 launch post reports FLEURS word error rate across 60 languages, an independent leaderboard position, and relative speed against competing transcription models. | Provider results Exact figures published in the official launch post | |||||
| MAI-Transcribe-1.5 launch evaluation Microsoft’s June 2, 2026 Foundry post reports the FLEURS word error rate improvement for MAI-Transcribe-1.5 over MAI-Transcribe-1. | Provider results Exact figures published in the official Foundry post | |||||
| MAI-Voice-2-Flash launch notes Microsoft AI’s July 23, 2026 post positions MAI-Voice-2-Flash as a faster, cheaper sibling of MAI-Voice-2 without publishing benchmark scores. | Provider results Relative claims only | |||||
| Sonic 3.6 launch evaluation Cartesia’s August 27, 2026 launch post reports independent controlled-voice Elo ratings, blind preference win rates against Eleven v3, response latency, and locale coverage for Sonic 3.6. | Provider results Exact figures published in the official launch post | |||||
| Realtime TTS-2 launch evaluation Inworld’s launch post reports an independent speech-arena ranking, time-to-first-audio, and language coverage for Realtime TTS-2. | Provider results Rank and latency published in the official launch post and docs; no Elo value given | |||||
| Realtime TTS-2 Flash launch notes Inworld’s model documentation reports time-to-first-byte and coverage for the Flash variant of Realtime TTS-2. | Provider results Documented figures only | |||||
| SeedRealtime launch notes ByteDance Seed’s August 5, 2026 announcement describes an end-to-end audio-visual full-duplex model and reports a halving of conversational pacing issues versus cascaded pipelines. | Provider results Relative claim only | |||||
| Hy ASR 3.0 preview launch notes Tencent’s Hunyuan site describes Hy ASR 3.0 preview as a Hy3-based recogniser with context-aware correction and broad dialect coverage; launch coverage cites word error rates near 3%. | Provider results Tencent’s site is descriptive; the WER values come from launch coverage and are marked secondary | |||||
| Deepgram Flux launch evaluation Deepgram’s October 2, 2025 launch announcement reports end-of-turn detection latency and Nova-3-level accuracy for Flux. | Provider results Latency figure published; accuracy stated relative to Nova-3 without a WER value | |||||
| Flux Multilingual launch evaluation Deepgram’s April 29, 2026 press release reports end-of-turn latency and language coverage for Flux Multilingual. | Provider results Latency and coverage published; accuracy stated as monolingual-grade without a WER value | |||||
| Pocket TTS project notes Kyutai’s TTS page and repository document Pocket TTS’s size, CPU real-time factor, and first-audio latency. | Provider results Documented figures only; no listening-test result | |||||
| Zonos 2 model evaluation Zyphra’s ZONOS2 GGUF model card reports word error rate, speaker similarity, and UTMOS for the reference F16 build. | Provider results Exact figures published on the official GGUF model card | |||||
| Granite Speech 5.0 TurboCTC model notes IBM’s model card and Granite 4.2 blog document the model’s size, training data, and throughput; Open ASR leaderboard results are published as images. | Provider results Throughput published in text; WER charts are images | |||||
| Cohere Transcribe Arabic model notes Cohere’s catalog and Hugging Face repository document the Arabic transcription model’s lineage and availability; no evaluation table is published. | Provider results No evaluation numbers published | |||||
| VibeVoice-ASR-Streaming model notes Microsoft’s model card documents streaming speaker-attributed transcription, hotword support, and ten-language coverage; evaluation results are published as an image. | Provider results Evaluation table published only as an image | |||||
| Qwen3-ASR model evaluation Qwen’s model card reports word error rates for Qwen3-ASR 1.7B and 0.6B across AMI, Earnings22, GigaSpeech, LibriSpeech, SPGISpeech, and VoxPopuli. | Provider results Exact figures published on the official model card | |||||
| AuK model notes Tencent’s model card documents AuK’s task coverage across speech generation, editing, enhancement, and separation; its benchmark chart is published as an image. | Provider results Benchmark chart published only as an image | |||||
| GPT-Live-1 API launch evaluation OpenAI’s API launch post reports spoken task success, full-duplex interactivity, tool calling, response quality, and turn-taking latency for GPT-Live-1 against GPT-Realtime-2.1 and GPT-Realtime-2. | Provider results Exact figures embedded in the official launch post charts | |||||
| Muse Voice Transcribe launch evaluation Meta’s launch evaluation reports streaming transcription accuracy, speaker diarization, and real-time audio-perception capabilities for Muse Voice Transcribe. | Provider results Exact figures published in the official launch post | |||||
| Artificial Analysis Speech-to-Speech Index A source-owned index for native audio models, with separate results for spoken reasoning, customer-service task completion, live preference, conversation flow, latency, and audio cost. | Live snapshot Published source table and repeatable page snapshot | |||||
| Big Bench Audio One thousand spoken reasoning questions test whether native audio models can answer correctly from audio and how quickly they begin replying. | Live snapshot Published source table and public audio dataset | |||||
| Speech Agent Arena People hold blind, paired voice conversations and choose the system they prefer; eligible tool-calling conversations also receive a task-success check. | Live snapshot Published Arena leaderboard with confidence intervals and sample counts | |||||
| Audio Realism Benchmark Blind listening comparisons measure how human text-to-speech output sounds across phone-agent, conversational, and explainer speech. | Live snapshot Machine-readable public leaderboard | |||||
| VoiceBench A spoken question-answering suite spanning open-ended, knowledge, reasoning, instruction-following, and safety tasks. | Live snapshot Machine-readable public leaderboard | |||||
| MMAU Measures expert-level understanding and reasoning across speech, environmental sound, and music. | Live snapshot Machine-readable benchmark-owner leaderboard | |||||
| MMAU-Pro Tests harder audio reasoning across speech, sound, music, spatial audio, multiple clips, voice chat, and instruction following. | Live snapshot Machine-readable benchmark-owner leaderboard | |||||
| AudioAgentBench Evaluates real-time voice agents on policy compliance, tool use, grounding, ambiguity, and state tracking. | Live snapshot Machine-readable public run records | |||||
| Full-Duplex-Bench v3 Tests full-duplex voice agents on multi-tool tasks with human speech, disfluencies, interruptions, and latency measurements. | Paper results Results transcribed from the official paper | |||||
| τ-Voice Extends τ-bench into spoken customer-service tasks to measure the capability gap between text and voice agents. | Live snapshot Machine-readable benchmark-owner submissions and reproducible evaluation code | |||||
| EVA-Bench Separates enterprise task accuracy from conversational experience across realistic service scenarios. | Live snapshot Machine-readable benchmark-owner leaderboard, code, and public dataset | |||||
| Voice Agent Latency Benchmark Measures the silence callers experience between finishing a turn and hearing a voice-agent platform begin its reply. | Live snapshot Machine-readable public API and reproducible raw-call artifacts | |||||
| VoiceAgentBench A multilingual voice-agent suite covering tools, workflows, multi-turn interaction, and safety. | Reference Paper and public dataset | |||||
| ADU-Bench Measures how voice systems handle ambiguous, disfluent, and underspecified spoken dialogue. | Reference Paper and benchmark resources |
Does the system understand and answer spoken requests?
Can it follow policy, use tools, and finish a workflow?
Can it handle turns, interruptions, ambiguity, and state?
Is the exchange natural, robust, and responsive?
How long do responses, tool calls, and full tasks take?