Artificial Analysis Speech-to-Speech Index
A source-owned index for native audio models, with separate results for spoken reasoning, customer-service task completion, live preference, conversation flow, latency, and audio cost.
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Benchmark profile
- Scope
- Native audio-input and audio-output models with all four index components; the source table also lists partial-result models
- Primary metric
- Equal-weighted source index (higher is better); component results stay separate
- Owner
- Artificial Analysis
- Available evidence
- Published source table and repeatable page snapshot
What it measures
How the protocol works
- 01
Native audio models receive spoken input and return audio output for reasoning and conversation checks.
- 02
Artificial Analysis runs customer-service tasks using its own τ-Voice setup; those results stay separate from Sierra’s submissions.
- 03
Human Arena preference and eligible task-completing tool calls provide the two live-conversation components.
- 04
Only models with reasoning, agentic, preference, and task-success results receive the equal-weighted index.
Available results
| Model | Index | Reasoning (rounded) | τ-Voice | Arena Elo (live) | Task success | Dynamics | First audio | Measured cost | Audio input price | Audio output price |
|---|---|---|---|---|---|---|---|---|---|---|
| Gemini 3.8 Live Extended Thinking (High) · Google | 82.6 | ~98% | 68.6% | 990 | 89.1% | 91.9% | 1.35s | $3.50/h | — | — |
| GPT-Live-1 (Astra, medium) · OpenAI | 81.5 | ~90% | 67.9% | 1048 | 87.4% | 94.9% | 1.34s | $5.83/h | — | — |
| Grok Voice Think Fast 2.0 High · SpaceXAI | 81.3 | ~97% | 56.5% | 1011 | 94.6% | 95.1% | 0.70s | $4.80/h | — | — |
| GPT-Live-1 (Sol, low) · OpenAI | 80.1 | ~89% | 59.3% | 1053 | 90.9% | 97.3% | 1.24s | $4.47/h | — | — |
| Gemini 3.8 Live · Google | 76.0 | ~92% | 30.1% | 1083 | 93.2% | 96.1% | 1.18s | $0.84/h | — | — |
| GPT-Realtime-2.1 High · OpenAI | 73.9 | ~96% | 45.7% | 928 | 91.5% | 95.7% | 1.21s | $10.75/h | — | — |
| GPT-Realtime-2 (High) · OpenAI | 73.6 | ~97% | 39.8% | 932 | 89.8% | 95.3% | 1.14s | $4.14/h | $1.15/h | $4.61/h |
| Grok Voice Think Fast 1.0 · SpaceXAI | 72.3 | ~97% | 52.1% | 908 | 80.7% | 77.8% | 1.25s | $3.00/h | — | — |
| Gemini 3.1 Flash Live High · Google | 71.5 | ~97% | 37.7% | 1063 | 71.8% | 74.3% | 2.99s | $1.75/h | $0.35/h | $1.38/h |
| GPT-Realtime-1.5 · OpenAI | 70.3 | ~81% | 38.8% | 1000 | 85.1% | 95.7% | 0.81s | $11.44/h | $1.15/h | $4.61/h |
| GPT-Realtime-2.1 Minimal · OpenAI | 70.3 | ~87% | 38.0% | 937 | 89.4% | 92.7% | 0.97s | $11.31/h | — | — |
| GPT Realtime (Aug '25) · OpenAI | 68.5 | ~83% | 30.4% | 974 | 89.4% | 93.9% | 0.98s | $11.08/h | $1.15/h | $4.61/h |
| Qwen Audio 3.0 Realtime Plus · Alibaba Cloud | 66.8 | ~99% | 54.6% | 775 | 77.8% | 98.4% | 1.54s | $4.42/h | $0.03/h | $0.18/h |
| Qwen Audio 3.0 Realtime Flash · Alibaba Cloud | 64.2 | ~96% | 35.9% | 850 | 81.7% | 96.9% | 1.55s | $4.77/h | — | — |
| Gemini 3.1 Flash Live Minimal · Google | 63.9 | ~71% | 26.2% | 1096 | 74.6% | 72.3% | 0.96s | $1.50/h | $0.35/h | $1.38/h |
| GPT-Realtime-2 (Minimal) · OpenAI | 62.7 | ~72% | 30.8% | 916 | 84.7% | 96.1% | 1.12s | $3.07/h | $1.15/h | $4.61/h |
| GPT Realtime Mini (Oct '25) · OpenAI | 56.8 | ~64% | 15.1% | 917 | 79.6% | 95.7% | 0.81s | $3.04/h | $0.36/h | $1.44/h |
| GPT-Realtime-2.1 Mini Minimal · OpenAI | 52.8 | ~63% | 22.5% | 840 | 76.7% | 91.8% | 0.85s | $4.60/h | — | — |