Reference coverage
ADU-Bench
Measures how voice systems handle ambiguous, disfluent, and underspecified spoken dialogue.
Source snapshots refreshed
Five measurement lanesVOICE / S2S
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Benchmark profile
- Scope
- 20,715 dialogues; 9 languages; 8,000+ real recordings
- Primary metric
- Skill- and ambiguity-level performance
- Owner
- ADU-Bench authors
- Available evidence
- Paper and benchmark resources
What it measures
Conversation dynamics
Can it handle turns, interruptions, ambiguity, and state?
Available results
No refreshable result table yet
The benchmark is tracked for coverage, but its current public results are not exposed here until model variants and protocol fields can be reproduced without ambiguity. Use the primary sources below for the published findings.
Interpretation limit
This is a diagnostic benchmark for ambiguity and disfluency, not an end-to-end service-agent ranking.