Skip to main content
Reference coverage

ADU-Bench

Measures how voice systems handle ambiguous, disfluent, and underspecified spoken dialogue.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
20,715 dialogues; 9 languages; 8,000+ real recordings
Primary metric
Skill- and ambiguity-level performance
Owner
ADU-Bench authors
Available evidence
Paper and benchmark resources

What it measures

Conversation dynamics
Can it handle turns, interruptions, ambiguity, and state?

Available results

No refreshable result table yet

The benchmark is tracked for coverage, but its current public results are not exposed here until model variants and protocol fields can be reproduced without ambiguity. Use the primary sources below for the published findings.

Interpretation limit

This is a diagnostic benchmark for ambiguity and disfluency, not an end-to-end service-agent ranking.