Skip to main content
Official paper results

Full-Duplex-Bench v3

Tests full-duplex voice agents on multi-tool tasks with human speech, disfluencies, interruptions, and latency measurements.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
100 examples; 79 scenarios; 12 speakers; 4 domains
Primary metric
Strict pass@1 and latency
Owner
Full-Duplex-Bench authors
Available evidence
Results transcribed from the official paper

What it measures

Task completion
Can it follow policy, use tools, and finish a workflow?
Conversation dynamics
Can it handle turns, interruptions, ambiguity, and state?
Latency
How long do responses, tool calls, and full tasks take?

Available results

Protocol details
ModelStrict pass@1First responseTool callTask completion
GPT Realtime 1.560%6.36s3.89s6.89s
Gemini Live 3.154%3.95s2.21s4.25s
Gemini Live 2.549%7.03s4.61s7.26s
Cascaded Whisper + GPT-4o + TTS45%8.78s3.15s10.12s
Grok Voice Agent43%5.97s0.63s6.65s
Ultravox Realtime41%3.88s6.01s8.40s

Interpretation limit

The six-system result table is a fixed paper snapshot, not a continuously updated leaderboard.