# ADU-Bench: Voice Benchmark Profile

> Measures how voice systems handle ambiguous, disfluent, and underspecified spoken dialogue.

- Scope: 20,715 dialogues; 9 languages; 8,000+ real recordings
- Measurement lanes: Conversation dynamics
- Primary metric: Skill- and ambiguity-level performance
- Available evidence: Paper and benchmark resources
- Owner: ADU-Bench authors
- Source snapshots refreshed: 2026-09-18

## Interpretation limit

This is a diagnostic benchmark for ambiguity and disfluency, not an end-to-end service-agent ranking.

## Primary sources

- [Paper](https://arxiv.org/abs/2412.05167)

Canonical page: https://benchlm.ai/voice-benchmarks/adu-bench
