Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes
Live source snapshot

τ-Voice

Extends τ-bench into spoken customer-service tasks to measure the capability gap between text and voice agents.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
278 tasks; clean and realistic acoustic conditions
Primary metric
Task success
Owner
Sierra Research
Available evidence
Machine-readable benchmark-owner submissions and reproducible evaluation code

What it measures

Task completion
Can it follow policy, use tools, and finish a workflow?
Conversation dynamics
Can it handle turns, interruptions, ambiguity, and state?

Available results

τ³-Voice task success by submitted voice system.
RankModelPass@1RetailAirlineTelecom
1gpt-live-1 · Backend: GPT-6 Astra (medium)81.7%79.0%82.0%84.2%
2Pine Voice Preview · enabled75.4%70.2%70.0%86.0%
3grok-voice-think-fast-1.0 · enabled67.3%62.3%66.0%73.7%
4grok-voice-think-fast-2.0 · high62.5%59.6%56.0%71.9%
5qwen3.5-omni-plus-realtime53.7%46.5%54.0%60.5%
6gemini-3.1-flash-live-preview-thinking-high · high43.9%45.6%64.0%21.9%
7gpt-realtime-2 · xhigh42.4%47.4%58.0%21.9%
8gpt-realtime-2 · minimal38.5%39.5%56.0%20.2%
9grok-voice-fast-1.038.3%38.6%36.0%40.4%
10gpt-realtime-1.535.3%44.7%40.0%21.1%
11Cascaded baseline31.2%28.9%48.0%16.7%
12gpt-realtime-1.030.4%36.0%36.0%19.3%
13gemini-3.1-flash-live-preview-thinking-minimal · minimal28.6%26.3%42.0%17.5%
14gemini-live-2.5-flash-native-audio25.8%29.8%30.0%17.5%

Interpretation limit

Commercial voice APIs and replacement voice IDs are required to reproduce the published setup.