Audio Realism Benchmark
The page imports Design Arena’s Bradley–Terry Elo-style rating, uncertainty, win rate, battle count, and average generation time. BenchLM does not refit the pairwise votes or merge the result with task benchmarks.
The method is intentionally conservative: show the benchmark owner’s values, keep incompatible protocols apart, and state where a table stops being current.
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Pull benchmark-owner leaderboards or run records when a stable public endpoint exists. Paper-only results remain fixed snapshots.
Keep each protocol’s metric names, scales, audio conditions, model variants, and judging method.
Do not mix reasoning, task success, conversation quality, experience, or latency into one weighted score.
Link the paper, code, dataset, and owner page so every displayed result can be traced to its source.
The page imports Design Arena’s Bradley–Terry Elo-style rating, uncertainty, win rate, battle count, and average generation time. BenchLM does not refit the pairwise votes or merge the result with task benchmarks.
The page transcribes model, architecture, access status, nine component metrics, and the benchmark owner’s overall score. It does not recalculate that overall value.
Public run-level checks are summed within model and audio condition. The displayed percentage is passed checks divided by eligible checks. Rehydrated runs are excluded from the human and synthetic tabs.
Strict pass@1 and mean latency values are transcribed from the official paper. This table changes only when the benchmark owner publishes a new result set.
τ-Voice, EVA-Bench, VoiceAgentBench, and ADU-Bench appear as coverage references until their result formats can be refreshed without model-name or protocol ambiguity.
Machine-readable public leaderboard. The leaderboard measures perceived human-likeness for the selected American-English voices and prompt mix. It does not measure factual accuracy, task completion, full-duplex interaction, or multilingual quality.
Machine-readable public leaderboard. Open-ended answers use an automatic model judge. Its overall score combines heterogeneous task scales using the benchmark owner’s method.
Machine-readable public run records. Scores are judge-derived checks aggregated across recent public runs. Human and synthesized audio remain separate views.
Results transcribed from the official paper. The six-system result table is a fixed paper snapshot, not a continuously updated leaderboard.
Paper and reproducible evaluation code. Commercial voice APIs and replacement voice IDs are required to reproduce the published setup.
Paper, code, and public dataset. Accuracy and experience are deliberately separate axes; neither should be collapsed into the other.
Paper and public dataset. Its broad multilingual coverage is primarily synthetic, so it should not stand in for human-audio robustness.