Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes
Live source snapshot

Speech Agent Arena

People hold blind, paired voice conversations and choose the system they prefer; eligible tool-calling conversations also receive a task-success check.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
35 scenarios: 15 with tool calls and 20 without; published model ranks, Elo intervals, sample counts, and eligible task-success intervals
Primary metric
Pairwise preference Elo and eligible task success rate (higher is better)
Owner
Artificial Analysis
Available evidence
Published Arena leaderboard with confidence intervals and sample counts

What it measures

Task completion
Can it follow policy, use tools, and finish a workflow?
Conversation dynamics
Can it handle turns, interruptions, ambiguity, and state?
Voice experience
Is the exchange natural, robust, and responsive?

How the protocol works

  1. 01

    A participant completes the same scenario with two hidden systems, then chooses an overall preference.

  2. 02

    Pairwise votes are fitted with Bradley–Terry maximum likelihood; GPT Realtime 1.5 anchors the Elo scale at 1000.

  3. 03

    Task success checks the correct final task-completing tool calls only for eligible conversations.

  4. 04

    The owner publishes rank ranges, 95% intervals, and sample counts with its leaderboard.

Available results

Speech Agent Arena owner ranks, preference Elo, sample counts, and eligible task success intervals.
RankSystemRank rangePreference EloElo 95% CISamplesTask successTSR 95% CI
1Gemini 3.1 Flash Live Minimal · Google1-21096-27/2784174.6%69.4%/79.2%
2Gemini 3.8 Live · Google1-41083-30/3048293.2%89.0%/95.8%
3Gemini 3.1 Flash Live High · Google2-51063-27/2781371.8%66.4%/76.7%
4GPT-Live-1 (Sol, low) · OpenAI3-61053-28/2865290.9%87.2%/93.7%
5GPT-Live-1 (Astra, medium) · OpenAI3-61048-28/2861087.4%82.9%/90.8%
6Cartesia Line (Default Cascaded System) · Cartesia3-91027-37/3719077.5%68.9%/84.3%
7Grok Voice Think Fast 2.0 High · SpaceXAI6-101011-28/2855294.6%91.4%/96.7%
8GPT-Realtime-1.5 · OpenAI810000/081185.1%80.8%/88.6%
9ElevenLabs Agents (Default Cascaded System) · ElevenLabs7-12993-26/2677390.5%86.7%/93.4%
10Gemini 3.8 Live Extended Thinking (High) · Google7-12990-27/2764389.1%84.7%/92.4%
11GPT Realtime (Aug '25) · OpenAI8-12974-26/2679789.4%85.2%/92.6%
12Nova 2.0 Sonic (Mar 2026) · Amazon Bedrock9-12974-26/2679557.1%51.1%/62.9%
13GPT-Realtime-2.1 Minimal · OpenAI13-19937-26/2679889.4%85.2%/92.5%
14GPT-Realtime-2 (High) · OpenAI13-20932-26/2676489.8%85.7%/92.9%
15GPT-Realtime-2.1 High · OpenAI13-20928-26/2674991.5%87.7%/94.3%
16GPT Realtime Mini (Oct '25) · OpenAI13-20917-26/2683079.6%74.6%/83.8%
17GPT-Realtime-2 (Minimal) · OpenAI13-20916-26/2683084.7%80.3%/88.3%
18Deepgram Voice Agent (Default Cascaded System) · Deepgram13-20914-26/2681773.7%68.5%/78.2%
19Nova Sonic · Amazon13-20913-26/2682757.1%51.6%/62.5%
20Grok Voice Think Fast 1.0 · SpaceXAI14-20908-26/2686480.7%75.8%/84.8%
21Inworld Realtime (Default Cascaded System) · Inworld21-24867-26/2681169.9%64.3%/75.0%
22Qwen3.5 Omni Plus Realtime · Alibaba Cloud21-25859-26/2682162.0%56.4%/67.2%
23Qwen Audio 3.0 Realtime Flash · Alibaba Cloud21-25850-40/4016681.7%70.1%/89.4%
24Qwen3.5 Omni Flash Realtime · Alibaba Cloud21-25841-26/2684729.1%24.0%/34.8%
25GPT-Realtime-2.1 Mini Minimal · OpenAI22-25840-27/2782176.7%71.1%/81.4%
26Qwen Audio 3.0 Realtime Plus · Alibaba Cloud26775-44/4417177.8%66.9%/85.8%

Interpretation limit

Preference and task success answer different questions. Task success excludes participant deviations and unverifiable calls; the table includes some provider-default cascaded systems alongside native audio models. These Arena scores remain separate from weighted text-model calibration and other voice benchmarks.