Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes
Official paper results

Full-Duplex-Bench v3

Tests full-duplex voice agents on multi-tool tasks with human speech, disfluencies, interruptions, and latency measurements.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
100 examples; 79 scenarios; 12 speakers; 4 domains
Primary metric
Strict pass@1 and latency
Owner
Full-Duplex-Bench authors
Available evidence
Results transcribed from the official paper

What it measures

Task completion
Can it follow policy, use tools, and finish a workflow?
Conversation dynamics
Can it handle turns, interruptions, ambiguity, and state?
Latency
How long do responses, tool calls, and full tasks take?

Available results

Protocol details
Full-duplex task results: strict pass rate and latency figures per model.
ModelStrict pass@1First responseTool callTask completion
GPT Realtime 1.560%6.36s3.89s6.89s
Gemini Live 3.154%3.95s2.21s4.25s
Gemini Live 2.549%7.03s4.61s7.26s
Cascaded Whisper + GPT-4o + TTS45%8.78s3.15s10.12s
Grok Voice Agent43%5.97s0.63s6.65s
Ultravox Realtime41%3.88s6.01s8.40s

Interpretation limit

The six-system result table is a fixed paper snapshot, not a continuously updated leaderboard.