Official paper results
Full-Duplex-Bench v3
Tests full-duplex voice agents on multi-tool tasks with human speech, disfluencies, interruptions, and latency measurements.
Source snapshots refreshed
Five measurement lanesVOICE / S2S
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Benchmark profile
- Scope
- 100 examples; 79 scenarios; 12 speakers; 4 domains
- Primary metric
- Strict pass@1 and latency
- Owner
- Full-Duplex-Bench authors
- Available evidence
- Results transcribed from the official paper
What it measures
Task completion
Can it follow policy, use tools, and finish a workflow?
Conversation dynamics
Can it handle turns, interruptions, ambiguity, and state?
Latency
How long do responses, tool calls, and full tasks take?
Available results
| Model | Strict pass@1 | First response | Tool call | Task completion |
|---|---|---|---|---|
| GPT Realtime 1.5 | 60% | 6.36s | 3.89s | 6.89s |
| Gemini Live 3.1 | 54% | 3.95s | 2.21s | 4.25s |
| Gemini Live 2.5 | 49% | 7.03s | 4.61s | 7.26s |
| Cascaded Whisper + GPT-4o + TTS | 45% | 8.78s | 3.15s | 10.12s |
| Grok Voice Agent | 43% | 5.97s | 0.63s | 6.65s |
| Ultravox Realtime | 41% | 3.88s | 6.01s | 8.40s |
Interpretation limit
The six-system result table is a fixed paper snapshot, not a continuously updated leaderboard.