Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes
Methodology

How the voice benchmark pages are built

The method is intentionally conservative: show the benchmark owner’s values, keep incompatible protocols apart, and state where a table stops being current.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Data handling

  1. 01Collect

    Pull benchmark-owner leaderboards or run records when a stable public endpoint exists. Paper-only results remain fixed snapshots.

  2. 02Preserve

    Keep each protocol’s metric names, scales, audio conditions, model variants, and judging method.

  3. 03Separate

    Preserve each source’s own aggregate while keeping unlike protocols out of weighted text-model scores.

  4. 04Expose

    Link the paper, code, dataset, and owner page so every displayed result can be traced to its source.

What is calculated

Artificial Analysis speech-to-speech results

The source’s published tables supply the index, Big Bench Audio accuracy, and Speech Agent Arena results. We preserve their displayed precision, intervals, and sample counts. The voice pages do not recalculate the index or use these rows in weighted text-model rankings.

Audio Realism Benchmark

The page imports Design Arena’s Bradley–Terry Elo-style rating, uncertainty, win rate, battle count, and average generation time. BenchLM does not refit the pairwise votes or merge the result with task benchmarks.

VoiceBench

The page transcribes model, architecture, access status, nine component metrics, and the benchmark owner’s overall score. It does not recalculate that overall value.

AudioAgentBench

Public run-level checks are summed within model and audio condition. The displayed percentage is passed checks divided by eligible checks. Rehydrated runs are excluded from the human and synthetic tabs.

Full-Duplex-Bench v3

Strict pass@1 and mean latency values are transcribed from the official paper. This table changes only when the benchmark owner publishes a new result set.

Reference benchmarks

τ-Voice, EVA-Bench, VoiceAgentBench, and ADU-Bench appear as coverage references until their result formats can be refreshed without model-name or protocol ambiguity.

Primary sources

Grok Voice Transcribe 2.0 launch evaluation

The short-phrase pair and the pricing are stated in the post’s prose. The four internal charts and the multilingual comparison render client-side, so their per-category values are not published in a readable form.. The internal sets are xAI’s own production traffic, not a released dataset, so nothing here is independently reproducible. The four per-category charts and the multilingual bar chart are drawn in the browser and carry no readable figures, so BenchLM stores only the values the post states in prose and does not read numbers off the plots. The leaderboard rank is an xAI claim about a third-party board that BenchLM does not track as a source; it is not a BenchLM measurement. These transcription results stay separate from BenchLM’s weighted text-model ranking.

Qwen3.8-Omni-Flash omni evaluation

Exact figures tabulated in the official launch post for all five compared systems. Qwen ran every system in this table itself and chose the harness for each agent row: WildClawBench-MM and AgenticVBench use Claude Code, UniClawBench uses OpenClaw, OmniGAIA uses no harness, and the static-versus-agent comparison uses Qwen Code. WildClawBench-MM covers only the multimodal subset of WildClawBench. Competitor runs use provider-specific media settings (Gemini 3.8 Flash at media_resolution=high, Seed 2.0 Lite at max_frame_tokens=384), so the comparison is not independently reproducible. These voice-protocol results stay separate from BenchLM’s weighted text-model ranking.

Canto launch evaluation

Exact figures labelled on the charts in the official launch post; the two dictation sets are private to Wispr. Wispr ran every system in this comparison itself, and the two headline sets are its own unreleased dictation data, so the results are not independently reproducible. Competitor runs used no contextual prompting, which is not how several of them are deployed, and the challenge-set winner (Gemini 3.1 Pro) is a frontier multimodal model Wispr describes as unsuitable for real-time dictation.

Gemini 3.8 Audio model evaluation

Exact bar-chart labels in Google’s model evaluation PDF, plus figures stated in the prose of the launch announcement; EVA-Bench scatterplot points are not tabulated. Google’s PDF combines evaluations run by separate teams using different APIs and effort settings. The Extended Thinking task rows use high effort; the EVA-Bench scatterplot labels are plotted without exact figures, and the launch announcement claims an EVA-Bench Pareto frontier without publishing a value. Google ran the Big Bench Audio and Speech Agent Arena figures itself and states them as standings against boards BenchLM refreshes from their owners, so a provider standing can disagree with the board snapshot. These voice-protocol results do not enter weighted text-model scores, and the plotted audio-hour costs are not canonical API prices.

GPT Audio 1.5 model documentation

Specifications only; OpenAI publishes no evaluation numbers for this model. This entry records what OpenAI documents for gpt-audio-1.5 rather than a benchmark result. BenchLM assigns no score.

GPT Transcribe model documentation

Specifications only. This entry records what OpenAI documents for gpt-transcribe rather than a benchmark result. BenchLM assigns no score.

GPT Live Transcribe model documentation

Specifications only. This entry records what OpenAI documents for gpt-live-transcribe rather than a benchmark result. BenchLM assigns no score.

Gemini 3.5 Transcribe launch evaluation

Exact figures published in the official launch post. These are provider-reported launch results against Google’s own prior model; the post gives no per-language breakdown or competitor rows.

MAI-Transcribe-2 launch evaluation

Exact figures published in the official launch post. These are provider-reported launch results; the post gives competitor rankings and speed ratios without their exact WER values.

MAI-Transcribe-1.5 launch evaluation

Exact figures published in the official Foundry post. Provider-reported results against Microsoft’s own previous model.

MAI-Voice-2-Flash launch notes

Relative claims only. Provider claims relative to Microsoft’s own MAI-Voice-2; no listening test or leaderboard result is published.

Sonic 3.6 launch evaluation

Exact figures published in the official launch post. Provider-reported results; the Elo figures are Cartesia’s citation of an independent speech leaderboard at launch and the preference tests are Cartesia-run.

Realtime TTS-2 launch evaluation

Rank and latency published in the official launch post and docs; no Elo value given. Provider-reported; Inworld cites the Speech Arena rank without the underlying Elo.

Realtime TTS-2 Flash launch notes

Documented figures only. Provider-documented latency; no arena or listening-test result is published for the Flash variant.

SeedRealtime launch notes

Relative claim only. Provider claim without a published benchmark table or competitor values.

Hy ASR 3.0 preview launch notes

Tencent’s site is descriptive; the WER values come from launch coverage and are marked secondary. The WER figures are reported by launch coverage (BigGo, August 5, 2026) rather than a Tencent page BenchLM could retrieve; treat them as secondary.

Deepgram Flux launch evaluation

Latency figure published; accuracy stated relative to Nova-3 without a WER value. Provider-reported launch figures; no independent WER table is published.

Flux Multilingual launch evaluation

Latency and coverage published; accuracy stated as monolingual-grade without a WER value. Provider-reported launch figures; no WER table is published.

Pocket TTS project notes

Documented figures only; no listening-test result. Provider-documented performance on reference hardware; no quality benchmark is published.

Zonos 2 model evaluation

Exact figures published on the official GGUF model card. Provider-run evaluation on Zyphra’s own set; the blog compares against other systems qualitatively only.

Granite Speech 5.0 TurboCTC model notes

Throughput published in text; WER charts are images. IBM publishes the Open ASR and FFASR leaderboard comparisons only as images, so no WER value is transcribed here.

Cohere Transcribe Arabic model notes

No evaluation numbers published. Cohere publishes no accuracy figures for this model at the time of tracking.

VibeVoice-ASR-Streaming model notes

Evaluation table published only as an image. No numeric results are transcribed because the card publishes them as an image.

Qwen3-ASR model evaluation

Exact figures published on the official model card. Provider-run evaluation on public English test sets.

AuK model notes

Benchmark chart published only as an image. No numeric results are transcribed because the card publishes them as an image.

GPT-Live-1 API launch evaluation

Exact figures embedded in the official launch post charts. These are provider-reported launch results against OpenAI’s own Realtime models, not an independent voice-agent ranking. The Tau3 and Tau Banking rows pair GPT-Live-1 with GPT-6 Astra at medium reasoning effort as the delegated backend, and the Full Duplex Bench v3 rows use the Terra backend at low effort. The Full Duplex Bench versions and runs differ from the v3 paper table BenchLM mirrors separately.

Muse Voice Transcribe launch evaluation

Exact figures published in the official launch post. These are provider-reported launch results, not an independent end-to-end voice-agent ranking. The post does not publish enough per-dataset detail to reconstruct either aggregate from raw runs.

Artificial Analysis Speech-to-Speech Index

Published source table and repeatable page snapshot. The index combines Big Bench Audio reasoning, Artificial Analysis’s τ-Voice implementation, frozen Arena preference, and Arena task success at 25% each. Some τ-Voice rows use fewer than three trials, and live Arena Elo can differ from the frozen value used in the index. The source-owned voice result does not enter weighted text-model scores; audio prices here are source observations, not canonical provider pricing.

Big Bench Audio

Published source table and public audio dataset. The audio is synthetic and English-only. A model judge marks spoken answers correct or incorrect. The published summary table rounds reasoning accuracy to whole percentages; the cost-per-hour figure comes from a fixed 40-question subset and is not a provider price.

Speech Agent Arena

Published Arena leaderboard with confidence intervals and sample counts. Preference and task success answer different questions. Task success excludes participant deviations and unverifiable calls; the table includes some provider-default cascaded systems alongside native audio models. These Arena scores remain separate from weighted text-model calibration and other voice benchmarks.

Audio Realism Benchmark

Machine-readable public leaderboard. The leaderboard measures perceived human-likeness for the selected American-English voices and prompt mix. It does not measure factual accuracy, task completion, full-duplex interaction, or multilingual quality.

VoiceBench

Machine-readable public leaderboard. Open-ended answers use an automatic model judge. Its overall score combines heterogeneous task scales using the benchmark owner’s method.

MMAU

Machine-readable benchmark-owner leaderboard. The refreshed table uses the parsed MMAU-v05.15.25 test results. It does not merge older benchmark versions or unverified community-reported rows.

MMAU-Pro

Machine-readable benchmark-owner leaderboard. The overall average spans heterogeneous audio tasks. Inspect speech, voice, spatial, multi-audio, and instruction-following columns before using it to choose a system.

AudioAgentBench

Machine-readable public run records. Scores are judge-derived checks aggregated across recent public runs. Human and synthesized audio remain separate views.

Full-Duplex-Bench v3

Results transcribed from the official paper. The six-system result table is a fixed paper snapshot, not a continuously updated leaderboard.

τ-Voice

Machine-readable benchmark-owner submissions and reproducible evaluation code. Commercial voice APIs and replacement voice IDs are required to reproduce the published setup.

EVA-Bench

Machine-readable benchmark-owner leaderboard, code, and public dataset. Accuracy and experience are deliberately separate axes; neither should be collapsed into the other.

Voice Agent Latency Benchmark

Machine-readable public API and reproducible raw-call artifacts. This board measures platform-level phone latency, not answer quality or raw speech-model latency. The carrier path and each platform’s deployed stack remain inside the measurement.

VoiceAgentBench

Paper and public dataset. Its broad multilingual coverage is primarily synthetic, so it should not stand in for human-audio robustness.

ADU-Bench

Paper and benchmark resources. This is a diagnostic benchmark for ambiguity and disfluency, not an end-to-end service-agent ranking.