Skip to main content
Methodology

How the voice benchmark pages are built

The method is intentionally conservative: show the benchmark owner’s values, keep incompatible protocols apart, and state where a table stops being current.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Data handling

  1. 01Collect

    Pull benchmark-owner leaderboards or run records when a stable public endpoint exists. Paper-only results remain fixed snapshots.

  2. 02Preserve

    Keep each protocol’s metric names, scales, audio conditions, model variants, and judging method.

  3. 03Separate

    Do not mix reasoning, task success, conversation quality, experience, or latency into one weighted score.

  4. 04Expose

    Link the paper, code, dataset, and owner page so every displayed result can be traced to its source.

What is calculated

Audio Realism Benchmark

The page imports Design Arena’s Bradley–Terry Elo-style rating, uncertainty, win rate, battle count, and average generation time. BenchLM does not refit the pairwise votes or merge the result with task benchmarks.

VoiceBench

The page transcribes model, architecture, access status, nine component metrics, and the benchmark owner’s overall score. It does not recalculate that overall value.

AudioAgentBench

Public run-level checks are summed within model and audio condition. The displayed percentage is passed checks divided by eligible checks. Rehydrated runs are excluded from the human and synthetic tabs.

Full-Duplex-Bench v3

Strict pass@1 and mean latency values are transcribed from the official paper. This table changes only when the benchmark owner publishes a new result set.

Reference benchmarks

τ-Voice, EVA-Bench, VoiceAgentBench, and ADU-Bench appear as coverage references until their result formats can be refreshed without model-name or protocol ambiguity.

Primary sources

Audio Realism Benchmark

Machine-readable public leaderboard. The leaderboard measures perceived human-likeness for the selected American-English voices and prompt mix. It does not measure factual accuracy, task completion, full-duplex interaction, or multilingual quality.

VoiceBench

Machine-readable public leaderboard. Open-ended answers use an automatic model judge. Its overall score combines heterogeneous task scales using the benchmark owner’s method.

AudioAgentBench

Machine-readable public run records. Scores are judge-derived checks aggregated across recent public runs. Human and synthesized audio remain separate views.

Full-Duplex-Bench v3

Results transcribed from the official paper. The six-system result table is a fixed paper snapshot, not a continuously updated leaderboard.

τ-Voice

Paper and reproducible evaluation code. Commercial voice APIs and replacement voice IDs are required to reproduce the published setup.

EVA-Bench

Paper, code, and public dataset. Accuracy and experience are deliberately separate axes; neither should be collapsed into the other.

VoiceAgentBench

Paper and public dataset. Its broad multilingual coverage is primarily synthetic, so it should not stand in for human-audio robustness.

ADU-Bench

Paper and benchmark resources. This is a diagnostic benchmark for ambiguity and disfluency, not an end-to-end service-agent ranking.