# Voice Benchmark Directory

> A coverage map for spoken reasoning, task completion, conversation dynamics, voice experience, and latency.

| Benchmark | Measurement lanes | Scope | Evidence |
|---|---|---|---|
| [Grok Voice Transcribe 2.0 launch evaluation](/voice-benchmarks/grok-voice-transcribe-2-launch-evaluation) ([md](/md/voice-benchmarks/grok-voice-transcribe-2-launch-evaluation.md)) | Voice experience | Four internal sets from production traffic — telephony (8 kHz customer-support calls, English), conversational (conversations with Grok, English), credentials (phone numbers, emails and addresses read aloud, English), and short phrases (voice-assistant utterances across 19 languages) — plus an independent public streaming speech-to-text leaderboard | The short-phrase pair and the pricing are stated in the post’s prose. The four internal charts and the multilingual comparison render client-side, so their per-category values are not published in a readable form. |
| [Qwen3.8-Omni-Flash omni evaluation](/voice-benchmarks/qwen3-8-omni-flash-evaluation) ([md](/md/voice-benchmarks/qwen3-8-omni-flash-evaluation.md)) | Spoken reasoning, Task completion, Conversation dynamics, Voice experience | Audio-visual agent benchmarks (WildClawBench-MM, UniClawBench, AgenticVBench, OmniGAIA), audio-visual understanding, reasoning, captioning and interaction sets (DailyOmni, WorldSense, AVUT, JoinAVBench, OmniVideoBench, Video-MME-v2, LVOmniBench, OmniCloze, OmniCap-IF, QIVD, StreamingBench), and audio sets covering multi-speaker ASR, multilingual ASR and speech translation, audio understanding and grounding, music understanding, and spoken interaction | Exact figures tabulated in the official launch post for all five compared systems |
| [Canto launch evaluation](/voice-benchmarks/wispr-canto-launch-evaluation) ([md](/md/voice-benchmarks/wispr-canto-launch-evaluation.md)) | Voice experience | 10 hours of opt-in Wispr Flow dictations from 2,300+ speakers, a 3-hour audio challenge set (noisy, low-volume, and short-utterance audio), and the English subsets of FLEURS, LibriSpeech, and Common Voice | Exact figures labelled on the charts in the official launch post; the two dictation sets are private to Wispr |
| [Gemini 3.8 Audio model evaluation](/voice-benchmarks/gemini-3-8-audio-evaluation) ([md](/md/voice-benchmarks/gemini-3-8-audio-evaluation.md)) | Task completion, Conversation dynamics, Voice experience | Live API voice-agent evaluations using ServiceNow EVA-Bench, a source-owned speech-to-speech index and τ-Voice, Sierra τ³-Banking, Big Bench Audio, and the Speech Agent Arena | Exact bar-chart labels in Google’s model evaluation PDF, plus figures stated in the prose of the launch announcement; EVA-Bench scatterplot points are not tabulated |
| [GPT Audio 1.5 model documentation](/voice-benchmarks/gpt-audio-1-5-launch-evaluation) ([md](/md/voice-benchmarks/gpt-audio-1-5-launch-evaluation.md)) | Voice experience, Latency | Modalities, context and output limits, endpoint support, and token pricing from the official model page | Specifications only; OpenAI publishes no evaluation numbers for this model |
| [GPT Transcribe model documentation](/voice-benchmarks/gpt-transcribe-launch-evaluation) ([md](/md/voice-benchmarks/gpt-transcribe-launch-evaluation.md)) | Voice experience, Latency | Endpoints, context features, and per-minute pricing from the official model page | Specifications only |
| [GPT Live Transcribe model documentation](/voice-benchmarks/gpt-live-transcribe-launch-evaluation) ([md](/md/voice-benchmarks/gpt-live-transcribe-launch-evaluation.md)) | Voice experience, Latency | Endpoint, latency controls, and per-minute pricing from the official model page | Specifications only |
| [Gemini 3.5 Transcribe launch evaluation](/voice-benchmarks/gemini-3-5-transcribe-launch-evaluation) ([md](/md/voice-benchmarks/gemini-3-5-transcribe-launch-evaluation.md)) | Voice experience, Latency | independent-leaderboard WER (non-streaming and streaming), FLEURS multilingual WER, and time-to-final-transcript versus Chirp 3 | Exact figures published in the official launch post |
| [MAI-Transcribe-2 launch evaluation](/voice-benchmarks/mai-transcribe-2-launch-evaluation) ([md](/md/voice-benchmarks/mai-transcribe-2-launch-evaluation.md)) | Voice experience, Latency | FLEURS average WER over 60 languages, independent WER leaderboard rank, and throughput comparisons versus GPT-Transcribe, Scribe v2, and Gemini 3.5 Transcribe | Exact figures published in the official launch post |
| [MAI-Transcribe-1.5 launch evaluation](/voice-benchmarks/mai-transcribe-1-5-launch-evaluation) ([md](/md/voice-benchmarks/mai-transcribe-1-5-launch-evaluation.md)) | Voice experience | FLEURS average WER over 25 languages and language coverage | Exact figures published in the official Foundry post |
| [MAI-Voice-2-Flash launch notes](/voice-benchmarks/mai-voice-2-flash-launch-evaluation) ([md](/md/voice-benchmarks/mai-voice-2-flash-launch-evaluation.md)) | Voice experience, Latency | Relative speed and cost versus MAI-Voice-2, per-character pricing, and deployment surfaces | Relative claims only |
| [Sonic 3.6 launch evaluation](/voice-benchmarks/cartesia-sonic-3-6-launch-evaluation) ([md](/md/voice-benchmarks/cartesia-sonic-3-6-launch-evaluation.md)) | Voice experience, Latency | independent controlled-voice Elo versus Sonic 3.5 and Eleven v3, blind head-to-head listener preference by locale, reply latency, and locale count | Exact figures published in the official launch post |
| [Realtime TTS-2 launch evaluation](/voice-benchmarks/inworld-tts-2-launch-evaluation) ([md](/md/voice-benchmarks/inworld-tts-2-launch-evaluation.md)) | Voice experience, Latency | independent speech-arena rank, P90 server-side time-to-first-byte, and language count | Rank and latency published in the official launch post and docs; no Elo value given |
| [Realtime TTS-2 Flash launch notes](/voice-benchmarks/inworld-tts-2-flash-launch-evaluation) ([md](/md/voice-benchmarks/inworld-tts-2-flash-launch-evaluation.md)) | Latency | Time-to-first-byte and language coverage from the official model docs | Documented figures only |
| [SeedRealtime launch notes](/voice-benchmarks/seedrealtime-launch-evaluation) ([md](/md/voice-benchmarks/seedrealtime-launch-evaluation.md)) | Conversation dynamics | Relative conversational-pacing claim versus cascaded ASR–LLM–TTS systems and architecture disclosures | Relative claim only |
| [Hy ASR 3.0 preview launch notes](/voice-benchmarks/hy-asr-3-0-preview-launch-evaluation) ([md](/md/voice-benchmarks/hy-asr-3-0-preview-launch-evaluation.md)) | Voice experience | Dialect coverage and context-aware recognition claims from Tencent, plus launch-coverage WER figures | Tencent’s site is descriptive; the WER values come from launch coverage and are marked secondary |
| [Deepgram Flux launch evaluation](/voice-benchmarks/deepgram-flux-launch-evaluation) ([md](/md/voice-benchmarks/deepgram-flux-launch-evaluation.md)) | Conversation dynamics, Latency | End-of-turn detection latency, accuracy relative to Nova-3, and GPU concurrency | Latency figure published; accuracy stated relative to Nova-3 without a WER value |
| [Flux Multilingual launch evaluation](/voice-benchmarks/deepgram-flux-multilingual-launch-evaluation) ([md](/md/voice-benchmarks/deepgram-flux-multilingual-launch-evaluation.md)) | Conversation dynamics, Latency | End-of-turn decision latency and supported languages | Latency and coverage published; accuracy stated as monolingual-grade without a WER value |
| [Pocket TTS project notes](/voice-benchmarks/kyutai-pocket-tts-launch-evaluation) ([md](/md/voice-benchmarks/kyutai-pocket-tts-launch-evaluation.md)) | Latency | Parameter count, CPU throughput, and time to first audio chunk | Documented figures only; no listening-test result |
| [Zonos 2 model evaluation](/voice-benchmarks/zonos-2-launch-evaluation) ([md](/md/voice-benchmarks/zonos-2-launch-evaluation.md)) | Voice experience | WER, speaker similarity, and UTMOS naturalness for the F16 reference build on Zyphra’s evaluation set | Exact figures published on the official GGUF model card |
| [Granite Speech 5.0 TurboCTC model notes](/voice-benchmarks/granite-speech-5-0-470m-turboctc-launch-evaluation) ([md](/md/voice-benchmarks/granite-speech-5-0-470m-turboctc-launch-evaluation.md)) | Latency | Parameter count, training hours, and RTFx throughput on a single H200 | Throughput published in text; WER charts are images |
| [Cohere Transcribe Arabic model notes](/voice-benchmarks/cohere-transcribe-arabic-07-2026-launch-evaluation) ([md](/md/voice-benchmarks/cohere-transcribe-arabic-07-2026-launch-evaluation.md)) | Voice experience | Base model, license, and availability | No evaluation numbers published |
| [VibeVoice-ASR-Streaming model notes](/voice-benchmarks/vibevoice-asr-streaming-7b-launch-evaluation) ([md](/md/voice-benchmarks/vibevoice-asr-streaming-7b-launch-evaluation.md)) | Conversation dynamics | Capabilities and language coverage from the model card | Evaluation table published only as an image |
| [Qwen3-ASR model evaluation](/voice-benchmarks/qwen3-asr-1-7b-launch-evaluation) ([md](/md/voice-benchmarks/qwen3-asr-1-7b-launch-evaluation.md)) | Voice experience | Mean WER over seven English test sets plus per-set values | Exact figures published on the official model card |
| [AuK model notes](/voice-benchmarks/tencent-auk-launch-evaluation) ([md](/md/voice-benchmarks/tencent-auk-launch-evaluation.md)) | Voice experience | Task coverage and model variants from the model card | Benchmark chart published only as an image |
| [GPT-Live-1 API launch evaluation](/voice-benchmarks/gpt-live-1-launch-evaluation) ([md](/md/voice-benchmarks/gpt-live-1-launch-evaluation.md)) | Task completion, Conversation dynamics, Latency | Tau3 (Voice) airline, retail, and telecom pass@1; Tau Banking (Voice) knowledge pass@1 over 97 tasks; Artificial Analysis Conversational Dynamics; Full Duplex Bench v1.5 interactivity, v1 turn-taking latency, and v3 tool calling and response quality | Exact figures embedded in the official launch post charts |
| [Muse Voice Transcribe launch evaluation](/voice-benchmarks/muse-voice-transcribe-evaluation) ([md](/md/voice-benchmarks/muse-voice-transcribe-evaluation.md)) | Conversation dynamics, Voice experience, Latency | Final-transcription streaming WER; average diarization error rate across AMI-IHM, AMI-SDM, and VoxConverse; multilingual and long-context capability disclosures | Exact figures published in the official launch post |
| [Artificial Analysis Speech-to-Speech Index](/voice-benchmarks/artificial-analysis-speech-to-speech-index) ([md](/md/voice-benchmarks/artificial-analysis-speech-to-speech-index.md)) | Spoken reasoning, Task completion, Conversation dynamics, Voice experience, Latency | Native audio-input and audio-output models with all four index components; the source table also lists partial-result models | Published source table and repeatable page snapshot |
| [Big Bench Audio](/voice-benchmarks/big-bench-audio) ([md](/md/voice-benchmarks/big-bench-audio.md)) | Spoken reasoning, Latency | 1,000 English audio questions; four Big Bench Hard-derived categories with 250 questions each; 23 synthetic voices | Published source table and public audio dataset |
| [Speech Agent Arena](/voice-benchmarks/speech-agent-arena) ([md](/md/voice-benchmarks/speech-agent-arena.md)) | Task completion, Conversation dynamics, Voice experience | 35 scenarios: 15 with tool calls and 20 without; published model ranks, Elo intervals, sample counts, and eligible task-success intervals | Published Arena leaderboard with confidence intervals and sample counts |
| [Audio Realism Benchmark](/voice-benchmarks/audio-realism-benchmark) ([md](/md/voice-benchmarks/audio-realism-benchmark.md)) | Voice experience | 500 held-out American-English prompts; female and male voices; human recordings included in the comparison pool | Machine-readable public leaderboard |
| [VoiceBench](/voice-benchmarks/voicebench) ([md](/md/voice-benchmarks/voicebench.md)) | Spoken reasoning | 11 dataset subsets; human and synthesized speech | Machine-readable public leaderboard |
| [MMAU](/voice-benchmarks/mmau) ([md](/md/voice-benchmarks/mmau.md)) | Spoken reasoning | 10,000 audio clips; 27 tasks; speech, sound, and music domains | Machine-readable benchmark-owner leaderboard |
| [MMAU-Pro](/voice-benchmarks/mmau-pro) ([md](/md/voice-benchmarks/mmau-pro.md)) | Spoken reasoning, Conversation dynamics | 5,305 human-curated questions; 49 skills; audio up to 10 minutes | Machine-readable benchmark-owner leaderboard |
| [AudioAgentBench](/voice-benchmarks/audioagentbench) ([md](/md/voice-benchmarks/audioagentbench.md)) | Task completion, Conversation dynamics | 6 service workflows; synthetic and human audio | Machine-readable public run records |
| [Full-Duplex-Bench v3](/voice-benchmarks/full-duplex-bench-v3) ([md](/md/voice-benchmarks/full-duplex-bench-v3.md)) | Task completion, Conversation dynamics, Latency | 100 examples; 79 scenarios; 12 speakers; 4 domains | Results transcribed from the official paper |
| [τ-Voice](/voice-benchmarks/tau-voice) ([md](/md/voice-benchmarks/tau-voice.md)) | Task completion, Conversation dynamics | 278 tasks; clean and realistic acoustic conditions | Machine-readable benchmark-owner submissions and reproducible evaluation code |
| [EVA-Bench](/voice-benchmarks/eva-bench) ([md](/md/voice-benchmarks/eva-bench.md)) | Task completion, Conversation dynamics, Voice experience | 213 scenarios; 3 enterprise domains; 12 systems | Machine-readable benchmark-owner leaderboard, code, and public dataset |
| [Voice Agent Latency Benchmark](/voice-benchmarks/voice-agent-latency) ([md](/md/voice-benchmarks/voice-agent-latency.md)) | Latency | Real synchronized phone calls with published recordings, per-turn measurements, discards, and configuration receipts | Machine-readable public API and reproducible raw-call artifacts |
| [VoiceAgentBench](/voice-benchmarks/voiceagentbench) ([md](/md/voice-benchmarks/voiceagentbench.md)) | Task completion, Conversation dynamics | 5,500+ synthetic queries; 7 languages | Paper and public dataset |
| [ADU-Bench](/voice-benchmarks/adu-bench) ([md](/md/voice-benchmarks/adu-bench.md)) | Conversation dynamics | 20,715 dialogues; 9 languages; 8,000+ real recordings | Paper and benchmark resources |

Canonical page: https://benchlm.ai/voice-benchmarks/benchmarks
