Skip to main content
Benchmark directory

A map of the voice benchmark landscape

Each benchmark covers a different slice of voice-system quality. The matrix makes those boundaries visible before you compare any scores.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

8 tracked benchmarks

BenchmarkSpoken reasoningTask completionConversation dynamicsVoice experienceLatencyEvidence
Audio Realism Benchmark

Blind listening comparisons measure how human text-to-speech output sounds across phone-agent, conversational, and explainer speech.

Live snapshot

Machine-readable public leaderboard

VoiceBench

A spoken question-answering suite spanning open-ended, knowledge, reasoning, instruction-following, and safety tasks.

Live snapshot

Machine-readable public leaderboard

AudioAgentBench

Evaluates real-time voice agents on policy compliance, tool use, grounding, ambiguity, and state tracking.

Live snapshot

Machine-readable public run records

Full-Duplex-Bench v3

Tests full-duplex voice agents on multi-tool tasks with human speech, disfluencies, interruptions, and latency measurements.

Paper results

Results transcribed from the official paper

τ-Voice

Extends τ-bench into spoken customer-service tasks to measure the capability gap between text and voice agents.

Reference

Paper and reproducible evaluation code

EVA-Bench

Separates enterprise task accuracy from conversational experience across realistic service scenarios.

Reference

Paper, code, and public dataset

VoiceAgentBench

A multilingual voice-agent suite covering tools, workflows, multi-turn interaction, and safety.

Reference

Paper and public dataset

ADU-Bench

Measures how voice systems handle ambiguous, disfluent, and underspecified spoken dialogue.

Reference

Paper and benchmark resources

Measurement lanes

01

Spoken reasoning

Does the system understand and answer spoken requests?

02

Task completion

Can it follow policy, use tools, and finish a workflow?

03

Conversation dynamics

Can it handle turns, interruptions, ambiguity, and state?

04

Voice experience

Is the exchange natural, robust, and responsive?

05

Latency

How long do responses, tool calls, and full tasks take?