Reference coverage
τ-Voice
Extends τ-bench into spoken customer-service tasks to measure the capability gap between text and voice agents.
Source snapshots refreshed
Five measurement lanesVOICE / S2S
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Benchmark profile
- Scope
- 278 tasks; clean and realistic acoustic conditions
- Primary metric
- Task success
- Owner
- Sierra Research
- Available evidence
- Paper and reproducible evaluation code
What it measures
Task completion
Can it follow policy, use tools, and finish a workflow?
Conversation dynamics
Can it handle turns, interruptions, ambiguity, and state?
Available results
No refreshable result table yet
The benchmark is tracked for coverage, but its current public results are not exposed here until model variants and protocol fields can be reproduced without ambiguity. Use the primary sources below for the published findings.
Interpretation limit
Commercial voice APIs and replacement voice IDs are required to reproduce the published setup.