Speech Agent Arena
People hold blind, paired voice conversations and choose the system they prefer; eligible tool-calling conversations also receive a task-success check.
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Benchmark profile
- Scope
- 35 scenarios: 15 with tool calls and 20 without; published model ranks, Elo intervals, sample counts, and eligible task-success intervals
- Primary metric
- Pairwise preference Elo and eligible task success rate (higher is better)
- Owner
- Artificial Analysis
- Available evidence
- Published Arena leaderboard with confidence intervals and sample counts
What it measures
How the protocol works
- 01
A participant completes the same scenario with two hidden systems, then chooses an overall preference.
- 02
Pairwise votes are fitted with Bradley–Terry maximum likelihood; GPT Realtime 1.5 anchors the Elo scale at 1000.
- 03
Task success checks the correct final task-completing tool calls only for eligible conversations.
- 04
The owner publishes rank ranges, 95% intervals, and sample counts with its leaderboard.
Available results
| Rank | System | Rank range | Preference Elo | Elo 95% CI | Samples | Task success | TSR 95% CI |
|---|---|---|---|---|---|---|---|
| 1 | Gemini 3.1 Flash Live Minimal · Google | 1-2 | 1096 | -27/27 | 841 | 74.6% | 69.4%/79.2% |
| 2 | Gemini 3.8 Live · Google | 1-4 | 1083 | -30/30 | 482 | 93.2% | 89.0%/95.8% |
| 3 | Gemini 3.1 Flash Live High · Google | 2-5 | 1063 | -27/27 | 813 | 71.8% | 66.4%/76.7% |
| 4 | GPT-Live-1 (Sol, low) · OpenAI | 3-6 | 1053 | -28/28 | 652 | 90.9% | 87.2%/93.7% |
| 5 | GPT-Live-1 (Astra, medium) · OpenAI | 3-6 | 1048 | -28/28 | 610 | 87.4% | 82.9%/90.8% |
| 6 | Cartesia Line (Default Cascaded System) · Cartesia | 3-9 | 1027 | -37/37 | 190 | 77.5% | 68.9%/84.3% |
| 7 | Grok Voice Think Fast 2.0 High · SpaceXAI | 6-10 | 1011 | -28/28 | 552 | 94.6% | 91.4%/96.7% |
| 8 | GPT-Realtime-1.5 · OpenAI | 8 | 1000 | 0/0 | 811 | 85.1% | 80.8%/88.6% |
| 9 | ElevenLabs Agents (Default Cascaded System) · ElevenLabs | 7-12 | 993 | -26/26 | 773 | 90.5% | 86.7%/93.4% |
| 10 | Gemini 3.8 Live Extended Thinking (High) · Google | 7-12 | 990 | -27/27 | 643 | 89.1% | 84.7%/92.4% |
| 11 | GPT Realtime (Aug '25) · OpenAI | 8-12 | 974 | -26/26 | 797 | 89.4% | 85.2%/92.6% |
| 12 | Nova 2.0 Sonic (Mar 2026) · Amazon Bedrock | 9-12 | 974 | -26/26 | 795 | 57.1% | 51.1%/62.9% |
| 13 | GPT-Realtime-2.1 Minimal · OpenAI | 13-19 | 937 | -26/26 | 798 | 89.4% | 85.2%/92.5% |
| 14 | GPT-Realtime-2 (High) · OpenAI | 13-20 | 932 | -26/26 | 764 | 89.8% | 85.7%/92.9% |
| 15 | GPT-Realtime-2.1 High · OpenAI | 13-20 | 928 | -26/26 | 749 | 91.5% | 87.7%/94.3% |
| 16 | GPT Realtime Mini (Oct '25) · OpenAI | 13-20 | 917 | -26/26 | 830 | 79.6% | 74.6%/83.8% |
| 17 | GPT-Realtime-2 (Minimal) · OpenAI | 13-20 | 916 | -26/26 | 830 | 84.7% | 80.3%/88.3% |
| 18 | Deepgram Voice Agent (Default Cascaded System) · Deepgram | 13-20 | 914 | -26/26 | 817 | 73.7% | 68.5%/78.2% |
| 19 | Nova Sonic · Amazon | 13-20 | 913 | -26/26 | 827 | 57.1% | 51.6%/62.5% |
| 20 | Grok Voice Think Fast 1.0 · SpaceXAI | 14-20 | 908 | -26/26 | 864 | 80.7% | 75.8%/84.8% |
| 21 | Inworld Realtime (Default Cascaded System) · Inworld | 21-24 | 867 | -26/26 | 811 | 69.9% | 64.3%/75.0% |
| 22 | Qwen3.5 Omni Plus Realtime · Alibaba Cloud | 21-25 | 859 | -26/26 | 821 | 62.0% | 56.4%/67.2% |
| 23 | Qwen Audio 3.0 Realtime Flash · Alibaba Cloud | 21-25 | 850 | -40/40 | 166 | 81.7% | 70.1%/89.4% |
| 24 | Qwen3.5 Omni Flash Realtime · Alibaba Cloud | 21-25 | 841 | -26/26 | 847 | 29.1% | 24.0%/34.8% |
| 25 | GPT-Realtime-2.1 Mini Minimal · OpenAI | 22-25 | 840 | -27/27 | 821 | 76.7% | 71.1%/81.4% |
| 26 | Qwen Audio 3.0 Realtime Plus · Alibaba Cloud | 26 | 775 | -44/44 | 171 | 77.8% | 66.9%/85.8% |