Grok Voice Transcribe 2.0 launch evaluationThe short-phrase pair and the pricing are stated in the post’s prose. The four internal charts and the multilingual comparison render client-side, so their per-category values are not published in a readable form.. The internal sets are xAI’s own production traffic, not a released dataset, so nothing here is independently reproducible. The four per-category charts and the multilingual bar chart are drawn in the browser and carry no readable figures, so BenchLM stores only the values the post states in prose and does not read numbers off the plots. The leaderboard rank is an xAI claim about a third-party board that BenchLM does not track as a source; it is not a BenchLM measurement. These transcription results stay separate from BenchLM’s weighted text-model ranking.
Qwen3.8-Omni-Flash omni evaluationExact figures tabulated in the official launch post for all five compared systems. Qwen ran every system in this table itself and chose the harness for each agent row: WildClawBench-MM and AgenticVBench use Claude Code, UniClawBench uses OpenClaw, OmniGAIA uses no harness, and the static-versus-agent comparison uses Qwen Code. WildClawBench-MM covers only the multimodal subset of WildClawBench. Competitor runs use provider-specific media settings (Gemini 3.8 Flash at media_resolution=high, Seed 2.0 Lite at max_frame_tokens=384), so the comparison is not independently reproducible. These voice-protocol results stay separate from BenchLM’s weighted text-model ranking.
Canto launch evaluationExact figures labelled on the charts in the official launch post; the two dictation sets are private to Wispr. Wispr ran every system in this comparison itself, and the two headline sets are its own unreleased dictation data, so the results are not independently reproducible. Competitor runs used no contextual prompting, which is not how several of them are deployed, and the challenge-set winner (Gemini 3.1 Pro) is a frontier multimodal model Wispr describes as unsuitable for real-time dictation.
Gemini 3.8 Audio model evaluationExact bar-chart labels in Google’s model evaluation PDF, plus figures stated in the prose of the launch announcement; EVA-Bench scatterplot points are not tabulated. Google’s PDF combines evaluations run by separate teams using different APIs and effort settings. The Extended Thinking task rows use high effort; the EVA-Bench scatterplot labels are plotted without exact figures, and the launch announcement claims an EVA-Bench Pareto frontier without publishing a value. Google ran the Big Bench Audio and Speech Agent Arena figures itself and states them as standings against boards BenchLM refreshes from their owners, so a provider standing can disagree with the board snapshot. These voice-protocol results do not enter weighted text-model scores, and the plotted audio-hour costs are not canonical API prices.
GPT Audio 1.5 model documentationSpecifications only; OpenAI publishes no evaluation numbers for this model. This entry records what OpenAI documents for gpt-audio-1.5 rather than a benchmark result. BenchLM assigns no score.
GPT Transcribe model documentationSpecifications only. This entry records what OpenAI documents for gpt-transcribe rather than a benchmark result. BenchLM assigns no score.
Gemini 3.5 Transcribe launch evaluationExact figures published in the official launch post. These are provider-reported launch results against Google’s own prior model; the post gives no per-language breakdown or competitor rows.
MAI-Transcribe-2 launch evaluationExact figures published in the official launch post. These are provider-reported launch results; the post gives competitor rankings and speed ratios without their exact WER values.
MAI-Voice-2-Flash launch notesRelative claims only. Provider claims relative to Microsoft’s own MAI-Voice-2; no listening test or leaderboard result is published.
Sonic 3.6 launch evaluationExact figures published in the official launch post. Provider-reported results; the Elo figures are Cartesia’s citation of an independent speech leaderboard at launch and the preference tests are Cartesia-run.
Realtime TTS-2 launch evaluationRank and latency published in the official launch post and docs; no Elo value given. Provider-reported; Inworld cites the Speech Arena rank without the underlying Elo.
Hy ASR 3.0 preview launch notesTencent’s site is descriptive; the WER values come from launch coverage and are marked secondary. The WER figures are reported by launch coverage (BigGo, August 5, 2026) rather than a Tencent page BenchLM could retrieve; treat them as secondary.
Deepgram Flux launch evaluationLatency figure published; accuracy stated relative to Nova-3 without a WER value. Provider-reported launch figures; no independent WER table is published.
Flux Multilingual launch evaluationLatency and coverage published; accuracy stated as monolingual-grade without a WER value. Provider-reported launch figures; no WER table is published.
Pocket TTS project notesDocumented figures only; no listening-test result. Provider-documented performance on reference hardware; no quality benchmark is published.
Zonos 2 model evaluationExact figures published on the official GGUF model card. Provider-run evaluation on Zyphra’s own set; the blog compares against other systems qualitatively only.
Granite Speech 5.0 TurboCTC model notesThroughput published in text; WER charts are images. IBM publishes the Open ASR and FFASR leaderboard comparisons only as images, so no WER value is transcribed here.
Qwen3-ASR model evaluationExact figures published on the official model card. Provider-run evaluation on public English test sets.
AuK model notesBenchmark chart published only as an image. No numeric results are transcribed because the card publishes them as an image.
GPT-Live-1 API launch evaluationExact figures embedded in the official launch post charts. These are provider-reported launch results against OpenAI’s own Realtime models, not an independent voice-agent ranking. The Tau3 and Tau Banking rows pair GPT-Live-1 with GPT-6 Astra at medium reasoning effort as the delegated backend, and the Full Duplex Bench v3 rows use the Terra backend at low effort. The Full Duplex Bench versions and runs differ from the v3 paper table BenchLM mirrors separately.
Muse Voice Transcribe launch evaluationExact figures published in the official launch post. These are provider-reported launch results, not an independent end-to-end voice-agent ranking. The post does not publish enough per-dataset detail to reconstruct either aggregate from raw runs.
Artificial Analysis Speech-to-Speech IndexPublished source table and repeatable page snapshot. The index combines Big Bench Audio reasoning, Artificial Analysis’s τ-Voice implementation, frozen Arena preference, and Arena task success at 25% each. Some τ-Voice rows use fewer than three trials, and live Arena Elo can differ from the frozen value used in the index. The source-owned voice result does not enter weighted text-model scores; audio prices here are source observations, not canonical provider pricing.
Big Bench AudioPublished source table and public audio dataset. The audio is synthetic and English-only. A model judge marks spoken answers correct or incorrect. The published summary table rounds reasoning accuracy to whole percentages; the cost-per-hour figure comes from a fixed 40-question subset and is not a provider price.
Speech Agent ArenaPublished Arena leaderboard with confidence intervals and sample counts. Preference and task success answer different questions. Task success excludes participant deviations and unverifiable calls; the table includes some provider-default cascaded systems alongside native audio models. These Arena scores remain separate from weighted text-model calibration and other voice benchmarks.
Audio Realism BenchmarkMachine-readable public leaderboard. The leaderboard measures perceived human-likeness for the selected American-English voices and prompt mix. It does not measure factual accuracy, task completion, full-duplex interaction, or multilingual quality.
VoiceBenchMachine-readable public leaderboard. Open-ended answers use an automatic model judge. Its overall score combines heterogeneous task scales using the benchmark owner’s method.
MMAUMachine-readable benchmark-owner leaderboard. The refreshed table uses the parsed MMAU-v05.15.25 test results. It does not merge older benchmark versions or unverified community-reported rows.
MMAU-ProMachine-readable benchmark-owner leaderboard. The overall average spans heterogeneous audio tasks. Inspect speech, voice, spatial, multi-audio, and instruction-following columns before using it to choose a system.
AudioAgentBenchMachine-readable public run records. Scores are judge-derived checks aggregated across recent public runs. Human and synthesized audio remain separate views.
Full-Duplex-Bench v3Results transcribed from the official paper. The six-system result table is a fixed paper snapshot, not a continuously updated leaderboard.
τ-VoiceMachine-readable benchmark-owner submissions and reproducible evaluation code. Commercial voice APIs and replacement voice IDs are required to reproduce the published setup.
EVA-BenchMachine-readable benchmark-owner leaderboard, code, and public dataset. Accuracy and experience are deliberately separate axes; neither should be collapsed into the other.
Voice Agent Latency BenchmarkMachine-readable public API and reproducible raw-call artifacts. This board measures platform-level phone latency, not answer quality or raw speech-model latency. The carrier path and each platform’s deployed stack remain inside the measurement.
VoiceAgentBenchPaper and public dataset. Its broad multilingual coverage is primarily synthetic, so it should not stand in for human-audio robustness.
ADU-BenchPaper and benchmark resources. This is a diagnostic benchmark for ambiguity and disfluency, not an end-to-end service-agent ranking.