Spoken reasoning
Does the system understand and answer spoken requests?
Each benchmark covers a different slice of voice-system quality. The matrix makes those boundaries visible before you compare any scores.
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
| Benchmark | Spoken reasoning | Task completion | Conversation dynamics | Voice experience | Latency | Evidence |
|---|---|---|---|---|---|---|
| Audio Realism Benchmark Blind listening comparisons measure how human text-to-speech output sounds across phone-agent, conversational, and explainer speech. | Live snapshot Machine-readable public leaderboard | |||||
| VoiceBench A spoken question-answering suite spanning open-ended, knowledge, reasoning, instruction-following, and safety tasks. | Live snapshot Machine-readable public leaderboard | |||||
| AudioAgentBench Evaluates real-time voice agents on policy compliance, tool use, grounding, ambiguity, and state tracking. | Live snapshot Machine-readable public run records | |||||
| Full-Duplex-Bench v3 Tests full-duplex voice agents on multi-tool tasks with human speech, disfluencies, interruptions, and latency measurements. | Paper results Results transcribed from the official paper | |||||
| τ-Voice Extends τ-bench into spoken customer-service tasks to measure the capability gap between text and voice agents. | Reference Paper and reproducible evaluation code | |||||
| EVA-Bench Separates enterprise task accuracy from conversational experience across realistic service scenarios. | Reference Paper, code, and public dataset | |||||
| VoiceAgentBench A multilingual voice-agent suite covering tools, workflows, multi-turn interaction, and safety. | Reference Paper and public dataset | |||||
| ADU-Bench Measures how voice systems handle ambiguous, disfluent, and underspecified spoken dialogue. | Reference Paper and benchmark resources |
Does the system understand and answer spoken requests?
Can it follow policy, use tools, and finish a workflow?
Can it handle turns, interruptions, ambiguity, and state?
Is the exchange natural, robust, and responsive?
How long do responses, tool calls, and full tasks take?