# Full-Duplex-Bench v3: Voice Benchmark Profile

> Tests full-duplex voice agents on multi-tool tasks with human speech, disfluencies, interruptions, and latency measurements.

- Scope: 100 examples; 79 scenarios; 12 speakers; 4 domains
- Measurement lanes: Task completion, Conversation dynamics, Latency
- Primary metric: Strict pass@1 and latency
- Available evidence: Results transcribed from the official paper
- Owner: Full-Duplex-Bench authors
- Source snapshots refreshed: 2026-09-18

## Interpretation limit

The six-system result table is a fixed paper snapshot, not a continuously updated leaderboard.

## Primary sources

- [Paper](https://arxiv.org/abs/2604.04847)
- [Code](https://github.com/DanielLin94144/Full-Duplex-Bench)

Canonical page: https://benchlm.ai/voice-benchmarks/full-duplex-bench-v3
