# τ-Voice: Voice Benchmark Profile

> Extends τ-bench into spoken customer-service tasks to measure the capability gap between text and voice agents.

- Scope: 278 tasks; clean and realistic acoustic conditions
- Measurement lanes: Task completion, Conversation dynamics
- Primary metric: Task success
- Available evidence: Machine-readable benchmark-owner submissions and reproducible evaluation code
- Owner: Sierra Research
- Source snapshots refreshed: 2026-09-18

## Interpretation limit

Commercial voice APIs and replacement voice IDs are required to reproduce the published setup.

## Primary sources

- [Paper](https://arxiv.org/abs/2603.13686)
- [Code](https://github.com/sierra-research/tau2-bench/tree/main/src/tau2/voice)
- [Dataset](https://sierra-tau-bench-public.s3.us-west-2.amazonaws.com/submissions/manifest.json)
- [Owner page](https://taubench.com/leaderboard?benchmark=voice)

Canonical page: https://benchlm.ai/voice-benchmarks/tau-voice
