# Voice Benchmark Methodology and Data Sources

> Results are pulled from benchmark owners, preserved in their original protocols, and never combined into a weighted voice score.

- Source snapshots refreshed: 2026-09-18
- VoiceBench: owner-provided model rows and overall scores are transcribed without recalculation.
- AudioAgentBench: passed and eligible checks are summed within each model and audio condition; rehydrated runs are excluded.
- τ³-Voice: Pass@1 stays separated by retail, airline, telecom, and banking domain from the owner submission manifest.
- EVA-Bench: task accuracy, conversational experience, confidence intervals, and response speed remain separate fields.
- MMAU and MMAU-Pro: benchmark-owner averages and task-family columns are preserved without merging benchmark versions.
- Voice Agent Latency: platform-level median and p95 TTFAB remain separate from model quality and raw speech-model latency.
- Full-Duplex-Bench v3: strict pass@1 and mean latency are transcribed from the official paper table.
- Other benchmarks remain references until their result formats can be refreshed without model or protocol ambiguity.

## Primary sources

- [Grok Voice Transcribe 2.0 launch evaluation](/voice-benchmarks/grok-voice-transcribe-2-launch-evaluation): [owner page](https://x.ai/news/grok-voice-transcribe-2). The internal sets are xAI’s own production traffic, not a released dataset, so nothing here is independently reproducible. The four per-category charts and the multilingual bar chart are drawn in the browser and carry no readable figures, so BenchLM stores only the values the post states in prose and does not read numbers off the plots. The leaderboard rank is an xAI claim about a third-party board that BenchLM does not track as a source; it is not a BenchLM measurement. These transcription results stay separate from BenchLM’s weighted text-model ranking.
- [Qwen3.8-Omni-Flash omni evaluation](/voice-benchmarks/qwen3-8-omni-flash-evaluation): [code](https://github.com/QwenLM/Qwen-MM-Plugins), [owner page](https://qwen.ai/blog?id=qwen3.8-omni-flash). Qwen ran every system in this table itself and chose the harness for each agent row: WildClawBench-MM and AgenticVBench use Claude Code, UniClawBench uses OpenClaw, OmniGAIA uses no harness, and the static-versus-agent comparison uses Qwen Code. WildClawBench-MM covers only the multimodal subset of WildClawBench. Competitor runs use provider-specific media settings (Gemini 3.8 Flash at media_resolution=high, Seed 2.0 Lite at max_frame_tokens=384), so the comparison is not independently reproducible. These voice-protocol results stay separate from BenchLM’s weighted text-model ranking.
- [Canto launch evaluation](/voice-benchmarks/wispr-canto-launch-evaluation): [owner page](https://wisprflow.ai/canto). Wispr ran every system in this comparison itself, and the two headline sets are its own unreleased dictation data, so the results are not independently reproducible. Competitor runs used no contextual prompting, which is not how several of them are deployed, and the challenge-set winner (Gemini 3.1 Pro) is a frontier multimodal model Wispr describes as unsuitable for real-time dictation.
- [Gemini 3.8 Audio model evaluation](/voice-benchmarks/gemini-3-8-audio-evaluation): [data](https://storage.googleapis.com/deepmind-media/gemini/gemini_3-8_live_model_evaluation.pdf), [owner page](https://deepmind.google/models/model-cards/gemini-3-8-audio/). Google’s PDF combines evaluations run by separate teams using different APIs and effort settings. The Extended Thinking task rows use high effort; the EVA-Bench scatterplot labels are plotted without exact figures, and the launch announcement claims an EVA-Bench Pareto frontier without publishing a value. Google ran the Big Bench Audio and Speech Agent Arena figures itself and states them as standings against boards BenchLM refreshes from their owners, so a provider standing can disagree with the board snapshot. These voice-protocol results do not enter weighted text-model scores, and the plotted audio-hour costs are not canonical API prices.
- [GPT Audio 1.5 model documentation](/voice-benchmarks/gpt-audio-1-5-launch-evaluation): [owner page](https://developers.openai.com/api/docs/models/gpt-audio-1.5). This entry records what OpenAI documents for gpt-audio-1.5 rather than a benchmark result. BenchLM assigns no score.
- [GPT Transcribe model documentation](/voice-benchmarks/gpt-transcribe-launch-evaluation): [owner page](https://developers.openai.com/api/docs/models/gpt-transcribe). This entry records what OpenAI documents for gpt-transcribe rather than a benchmark result. BenchLM assigns no score.
- [GPT Live Transcribe model documentation](/voice-benchmarks/gpt-live-transcribe-launch-evaluation): [owner page](https://developers.openai.com/api/docs/models/gpt-live-transcribe). This entry records what OpenAI documents for gpt-live-transcribe rather than a benchmark result. BenchLM assigns no score.
- [Gemini 3.5 Transcribe launch evaluation](/voice-benchmarks/gemini-3-5-transcribe-launch-evaluation): [owner page](https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-5-transcribe/). These are provider-reported launch results against Google’s own prior model; the post gives no per-language breakdown or competitor rows.
- [MAI-Transcribe-2 launch evaluation](/voice-benchmarks/mai-transcribe-2-launch-evaluation): [owner page](https://microsoft.ai/news/mai-transcribe-2-is-the-fastest-most-accurate-and-cheapest-speech-recognition-model-in-the-world/). These are provider-reported launch results; the post gives competitor rankings and speed ratios without their exact WER values.
- [MAI-Transcribe-1.5 launch evaluation](/voice-benchmarks/mai-transcribe-1-5-launch-evaluation): [owner page](https://techcommunity.microsoft.com/blog/azure-ai-foundry-blog/new-mai-models-in-microsoft-foundry-across-text-image-voice-and-speech/4524632). Provider-reported results against Microsoft’s own previous model.
- [MAI-Voice-2-Flash launch notes](/voice-benchmarks/mai-voice-2-flash-launch-evaluation): [owner page](https://microsoft.ai/news/introducing-mai-image-2-5-pro-and-mai-voice-2-flash/). Provider claims relative to Microsoft’s own MAI-Voice-2; no listening test or leaderboard result is published.
- [Sonic 3.6 launch evaluation](/voice-benchmarks/cartesia-sonic-3-6-launch-evaluation): [owner page](https://www.cartesia.ai/blog/sonic-3.6). Provider-reported results; the Elo figures are Cartesia’s citation of an independent speech leaderboard at launch and the preference tests are Cartesia-run.
- [Realtime TTS-2 launch evaluation](/voice-benchmarks/inworld-tts-2-launch-evaluation): [owner page](https://inworld.ai/blog/realtime-tts-2). Provider-reported; Inworld cites the Speech Arena rank without the underlying Elo.
- [Realtime TTS-2 Flash launch notes](/voice-benchmarks/inworld-tts-2-flash-launch-evaluation): [owner page](https://docs.inworld.ai/tts/tts-models). Provider-documented latency; no arena or listening-test result is published for the Flash variant.
- [SeedRealtime launch notes](/voice-benchmarks/seedrealtime-launch-evaluation): [owner page](https://seed.bytedance.com/en/SeedRealtime). Provider claim without a published benchmark table or competitor values.
- [Hy ASR 3.0 preview launch notes](/voice-benchmarks/hy-asr-3-0-preview-launch-evaluation): [owner page](https://hunyuan.tencent.com/). The WER figures are reported by launch coverage (BigGo, August 5, 2026) rather than a Tencent page BenchLM could retrieve; treat them as secondary.
- [Deepgram Flux launch evaluation](/voice-benchmarks/deepgram-flux-launch-evaluation): [owner page](https://finance.yahoo.com/news/deepgram-launches-flux-world-first-183000065.html). Provider-reported launch figures; no independent WER table is published.
- [Flux Multilingual launch evaluation](/voice-benchmarks/deepgram-flux-multilingual-launch-evaluation): [owner page](https://deepgram.com/learn/deepgram-launches-flux-multilingual-press-release). Provider-reported launch figures; no WER table is published.
- [Pocket TTS project notes](/voice-benchmarks/kyutai-pocket-tts-launch-evaluation): [owner page](https://github.com/kyutai-labs/pocket-tts). Provider-documented performance on reference hardware; no quality benchmark is published.
- [Zonos 2 model evaluation](/voice-benchmarks/zonos-2-launch-evaluation): [owner page](https://huggingface.co/Zyphra/ZONOS2-GGUF). Provider-run evaluation on Zyphra’s own set; the blog compares against other systems qualitatively only.
- [Granite Speech 5.0 TurboCTC model notes](/voice-benchmarks/granite-speech-5-0-470m-turboctc-launch-evaluation): [owner page](https://huggingface.co/ibm-granite/granite-speech-5.0-470m-turboctc). IBM publishes the Open ASR and FFASR leaderboard comparisons only as images, so no WER value is transcribed here.
- [Cohere Transcribe Arabic model notes](/voice-benchmarks/cohere-transcribe-arabic-07-2026-launch-evaluation): [owner page](https://docs.cohere.com/docs/models). Cohere publishes no accuracy figures for this model at the time of tracking.
- [VibeVoice-ASR-Streaming model notes](/voice-benchmarks/vibevoice-asr-streaming-7b-launch-evaluation): [owner page](https://huggingface.co/microsoft/VibeVoice-ASR-Streaming-7B). No numeric results are transcribed because the card publishes them as an image.
- [Qwen3-ASR model evaluation](/voice-benchmarks/qwen3-asr-1-7b-launch-evaluation): [owner page](https://huggingface.co/Qwen/Qwen3-ASR-1.7B-hf). Provider-run evaluation on public English test sets.
- [AuK model notes](/voice-benchmarks/tencent-auk-launch-evaluation): [owner page](https://huggingface.co/tencent/AuK). No numeric results are transcribed because the card publishes them as an image.
- [GPT-Live-1 API launch evaluation](/voice-benchmarks/gpt-live-1-launch-evaluation): [owner page](https://openai.com/index/introducing-gpt-live-1-in-the-api/). These are provider-reported launch results against OpenAI’s own Realtime models, not an independent voice-agent ranking. The Tau3 and Tau Banking rows pair GPT-Live-1 with GPT-6 Astra at medium reasoning effort as the delegated backend, and the Full Duplex Bench v3 rows use the Terra backend at low effort. The Full Duplex Bench versions and runs differ from the v3 paper table BenchLM mirrors separately.
- [Muse Voice Transcribe launch evaluation](/voice-benchmarks/muse-voice-transcribe-evaluation): [owner page](https://research.meta.ai/blog/introducing-muse-voice-transcribe). These are provider-reported launch results, not an independent end-to-end voice-agent ranking. The post does not publish enough per-dataset detail to reconstruct either aggregate from raw runs.
- [Artificial Analysis Speech-to-Speech Index](/voice-benchmarks/artificial-analysis-speech-to-speech-index): [data](https://huggingface.co/datasets/ArtificialAnalysis/big_bench_audio), [owner page](https://artificialanalysis.ai/speech-to-speech). The index combines Big Bench Audio reasoning, Artificial Analysis’s τ-Voice implementation, frozen Arena preference, and Arena task success at 25% each. Some τ-Voice rows use fewer than three trials, and live Arena Elo can differ from the frozen value used in the index. The source-owned voice result does not enter weighted text-model scores; audio prices here are source observations, not canonical provider pricing.
- [Big Bench Audio](/voice-benchmarks/big-bench-audio): [data](https://huggingface.co/datasets/ArtificialAnalysis/big_bench_audio), [owner page](https://artificialanalysis.ai/speech-to-speech). The audio is synthetic and English-only. A model judge marks spoken answers correct or incorrect. The published summary table rounds reasoning accuracy to whole percentages; the cost-per-hour figure comes from a fixed 40-question subset and is not a provider price.
- [Speech Agent Arena](/voice-benchmarks/speech-agent-arena): [owner page](https://artificialanalysis.ai/speech-to-speech/arena). Preference and task success answer different questions. Task success excludes participant deviations and unverifiable calls; the table includes some provider-default cascaded systems alongside native audio models. These Arena scores remain separate from weighted text-model calibration and other voice benchmarks.
- [Audio Realism Benchmark](/voice-benchmarks/audio-realism-benchmark): [owner page](https://www.designarena.ai/methodology/audio-realism-benchmark). The leaderboard measures perceived human-likeness for the selected American-English voices and prompt mix. It does not measure factual accuracy, task completion, full-duplex interaction, or multilingual quality.
- [VoiceBench](/voice-benchmarks/voicebench): [paper](https://arxiv.org/abs/2410.17196), [code](https://github.com/MatthewCYM/VoiceBench), [data](https://huggingface.co/datasets/VoiceBench/VoiceBench), [owner page](https://matthewcym.github.io/VoiceBench/). Open-ended answers use an automatic model judge. Its overall score combines heterogeneous task scales using the benchmark owner’s method.
- [MMAU](/voice-benchmarks/mmau): [paper](https://arxiv.org/abs/2410.19168), [code](https://github.com/Sakshi113/MMAU), [data](https://huggingface.co/datasets/gamma-lab-umd/MMAU-test), [owner page](https://sakshi113.github.io/mmau_homepage/#leaderboard-v15-parsed). The refreshed table uses the parsed MMAU-v05.15.25 test results. It does not merge older benchmark versions or unverified community-reported rows.
- [MMAU-Pro](/voice-benchmarks/mmau-pro): [paper](https://arxiv.org/abs/2508.13992), [code](https://github.com/sonalkum/MMAUPro), [data](https://huggingface.co/datasets/gamma-lab-umd/MMAU-Pro), [owner page](https://sonalkum.github.io/mmau-pro/#results). The overall average spans heterogeneous audio tasks. Inspect speech, voice, spatial, multi-audio, and instruction-following columns before using it to choose a system.
- [AudioAgentBench](/voice-benchmarks/audioagentbench): [code](https://github.com/Design-Arena/audio-agent-bench), [owner page](https://audioarena.ai/methodology). Scores are judge-derived checks aggregated across recent public runs. Human and synthesized audio remain separate views.
- [Full-Duplex-Bench v3](/voice-benchmarks/full-duplex-bench-v3): [paper](https://arxiv.org/abs/2604.04847), [code](https://github.com/DanielLin94144/Full-Duplex-Bench). The six-system result table is a fixed paper snapshot, not a continuously updated leaderboard.
- [τ-Voice](/voice-benchmarks/tau-voice): [paper](https://arxiv.org/abs/2603.13686), [code](https://github.com/sierra-research/tau2-bench/tree/main/src/tau2/voice), [data](https://sierra-tau-bench-public.s3.us-west-2.amazonaws.com/submissions/manifest.json), [owner page](https://taubench.com/leaderboard?benchmark=voice). Commercial voice APIs and replacement voice IDs are required to reproduce the published setup.
- [EVA-Bench](/voice-benchmarks/eva-bench): [paper](https://arxiv.org/abs/2605.13841), [code](https://github.com/ServiceNow/eva), [data](https://huggingface.co/datasets/ServiceNow-AI/eva), [owner page](https://servicenow.github.io/eva). Accuracy and experience are deliberately separate axes; neither should be collapsed into the other.
- [Voice Agent Latency Benchmark](/voice-benchmarks/voice-agent-latency): [code](https://github.com/openbenchmarks-labs/voice-agent-latency), [data](https://openbenchmarks.com/api/benchmarks/voice-agent-latency), [owner page](https://openbenchmarks.com/voice-agent-latency). This board measures platform-level phone latency, not answer quality or raw speech-model latency. The carrier path and each platform’s deployed stack remain inside the measurement.
- [VoiceAgentBench](/voice-benchmarks/voiceagentbench): [paper](https://arxiv.org/abs/2510.07978), [data](https://huggingface.co/datasets/krutrim-ai-labs/VoiceAgentBench). Its broad multilingual coverage is primarily synthetic, so it should not stand in for human-audio robustness.
- [ADU-Bench](/voice-benchmarks/adu-bench): [paper](https://arxiv.org/abs/2412.05167). This is a diagnostic benchmark for ambiguity and disfluency, not an end-to-end service-agent ranking.

Canonical page: https://benchlm.ai/voice-benchmarks/methodology
