Skip to main content
Reference coverage

τ-Voice

Extends τ-bench into spoken customer-service tasks to measure the capability gap between text and voice agents.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
278 tasks; clean and realistic acoustic conditions
Primary metric
Task success
Owner
Sierra Research
Available evidence
Paper and reproducible evaluation code

What it measures

Task completion
Can it follow policy, use tools, and finish a workflow?
Conversation dynamics
Can it handle turns, interruptions, ambiguity, and state?

Available results

No refreshable result table yet

The benchmark is tracked for coverage, but its current public results are not exposed here until model variants and protocol fields can be reproduced without ambiguity. Use the primary sources below for the published findings.

Interpretation limit

Commercial voice APIs and replacement voice IDs are required to reproduce the published setup.