# Speech Agent Arena: Voice Benchmark Profile

> People hold blind, paired voice conversations and choose the system they prefer; eligible tool-calling conversations also receive a task-success check.

- Scope: 35 scenarios: 15 with tool calls and 20 without; published model ranks, Elo intervals, sample counts, and eligible task-success intervals
- Measurement lanes: Task completion, Conversation dynamics, Voice experience
- Primary metric: Pairwise preference Elo and eligible task success rate (higher is better)
- Available evidence: Published Arena leaderboard with confidence intervals and sample counts
- Owner: Artificial Analysis
- Source snapshots refreshed: 2026-09-18

## Interpretation limit

Preference and task success answer different questions. Task success excludes participant deviations and unverifiable calls; the table includes some provider-default cascaded systems alongside native audio models. These Arena scores remain separate from weighted text-model calibration and other voice benchmarks.

## Primary sources

- [Owner page](https://artificialanalysis.ai/speech-to-speech/arena)

Canonical page: https://benchlm.ai/voice-benchmarks/speech-agent-arena
