Live source snapshot
τ-Voice
Extends τ-bench into spoken customer-service tasks to measure the capability gap between text and voice agents.
Source snapshots refreshed
Five measurement lanesVOICE / S2S
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Benchmark profile
- Scope
- 278 tasks; clean and realistic acoustic conditions
- Primary metric
- Task success
- Owner
- Sierra Research
- Available evidence
- Machine-readable benchmark-owner submissions and reproducible evaluation code
What it measures
Task completion
Can it follow policy, use tools, and finish a workflow?
Conversation dynamics
Can it handle turns, interruptions, ambiguity, and state?
Available results
| Rank | Model | Pass@1 | Retail | Airline | Telecom |
|---|---|---|---|---|---|
| 1 | gpt-live-1 · Backend: GPT-6 Astra (medium) | 81.7% | 79.0% | 82.0% | 84.2% |
| 2 | Pine Voice Preview · enabled | 75.4% | 70.2% | 70.0% | 86.0% |
| 3 | grok-voice-think-fast-1.0 · enabled | 67.3% | 62.3% | 66.0% | 73.7% |
| 4 | grok-voice-think-fast-2.0 · high | 62.5% | 59.6% | 56.0% | 71.9% |
| 5 | qwen3.5-omni-plus-realtime | 53.7% | 46.5% | 54.0% | 60.5% |
| 6 | gemini-3.1-flash-live-preview-thinking-high · high | 43.9% | 45.6% | 64.0% | 21.9% |
| 7 | gpt-realtime-2 · xhigh | 42.4% | 47.4% | 58.0% | 21.9% |
| 8 | gpt-realtime-2 · minimal | 38.5% | 39.5% | 56.0% | 20.2% |
| 9 | grok-voice-fast-1.0 | 38.3% | 38.6% | 36.0% | 40.4% |
| 10 | gpt-realtime-1.5 | 35.3% | 44.7% | 40.0% | 21.1% |
| 11 | Cascaded baseline | 31.2% | 28.9% | 48.0% | 16.7% |
| 12 | gpt-realtime-1.0 | 30.4% | 36.0% | 36.0% | 19.3% |
| 13 | gemini-3.1-flash-live-preview-thinking-minimal · minimal | 28.6% | 26.3% | 42.0% | 17.5% |
| 14 | gemini-live-2.5-flash-native-audio | 25.8% | 29.8% | 30.0% | 17.5% |
Interpretation limit
Commercial voice APIs and replacement voice IDs are required to reproduce the published setup.