Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes
Live source snapshot

Artificial Analysis Speech-to-Speech Index

A source-owned index for native audio models, with separate results for spoken reasoning, customer-service task completion, live preference, conversation flow, latency, and audio cost.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
Native audio-input and audio-output models with all four index components; the source table also lists partial-result models
Primary metric
Equal-weighted source index (higher is better); component results stay separate
Owner
Artificial Analysis
Available evidence
Published source table and repeatable page snapshot

What it measures

Spoken reasoning
Does the system understand and answer spoken requests?
Task completion
Can it follow policy, use tools, and finish a workflow?
Conversation dynamics
Can it handle turns, interruptions, ambiguity, and state?
Voice experience
Is the exchange natural, robust, and responsive?
Latency
How long do responses, tool calls, and full tasks take?

How the protocol works

  1. 01

    Native audio models receive spoken input and return audio output for reasoning and conversation checks.

  2. 02

    Artificial Analysis runs customer-service tasks using its own τ-Voice setup; those results stay separate from Sierra’s submissions.

  3. 03

    Human Arena preference and eligible task-completing tool calls provide the two live-conversation components.

  4. 04

    Only models with reasoning, agentic, preference, and task-success results receive the equal-weighted index.

Available results

Artificial Analysis speech-to-speech index and its separate source components. Published table precision.
ModelIndexReasoning (rounded)τ-VoiceArena Elo (live)Task successDynamicsFirst audioMeasured costAudio input priceAudio output price
Gemini 3.8 Live Extended Thinking (High) · Google82.6~98%68.6%99089.1%91.9%1.35s$3.50/h
GPT-Live-1 (Astra, medium) · OpenAI81.5~90%67.9%104887.4%94.9%1.34s$5.83/h
Grok Voice Think Fast 2.0 High · SpaceXAI81.3~97%56.5%101194.6%95.1%0.70s$4.80/h
GPT-Live-1 (Sol, low) · OpenAI80.1~89%59.3%105390.9%97.3%1.24s$4.47/h
Gemini 3.8 Live · Google76.0~92%30.1%108393.2%96.1%1.18s$0.84/h
GPT-Realtime-2.1 High · OpenAI73.9~96%45.7%92891.5%95.7%1.21s$10.75/h
GPT-Realtime-2 (High) · OpenAI73.6~97%39.8%93289.8%95.3%1.14s$4.14/h$1.15/h$4.61/h
Grok Voice Think Fast 1.0 · SpaceXAI72.3~97%52.1%90880.7%77.8%1.25s$3.00/h
Gemini 3.1 Flash Live High · Google71.5~97%37.7%106371.8%74.3%2.99s$1.75/h$0.35/h$1.38/h
GPT-Realtime-1.5 · OpenAI70.3~81%38.8%100085.1%95.7%0.81s$11.44/h$1.15/h$4.61/h
GPT-Realtime-2.1 Minimal · OpenAI70.3~87%38.0%93789.4%92.7%0.97s$11.31/h
GPT Realtime (Aug '25) · OpenAI68.5~83%30.4%97489.4%93.9%0.98s$11.08/h$1.15/h$4.61/h
Qwen Audio 3.0 Realtime Plus · Alibaba Cloud66.8~99%54.6%77577.8%98.4%1.54s$4.42/h$0.03/h$0.18/h
Qwen Audio 3.0 Realtime Flash · Alibaba Cloud64.2~96%35.9%85081.7%96.9%1.55s$4.77/h
Gemini 3.1 Flash Live Minimal · Google63.9~71%26.2%109674.6%72.3%0.96s$1.50/h$0.35/h$1.38/h
GPT-Realtime-2 (Minimal) · OpenAI62.7~72%30.8%91684.7%96.1%1.12s$3.07/h$1.15/h$4.61/h
GPT Realtime Mini (Oct '25) · OpenAI56.8~64%15.1%91779.6%95.7%0.81s$3.04/h$0.36/h$1.44/h
GPT-Realtime-2.1 Mini Minimal · OpenAI52.8~63%22.5%84076.7%91.8%0.85s$4.60/h

Interpretation limit

The index combines Big Bench Audio reasoning, Artificial Analysis’s τ-Voice implementation, frozen Arena preference, and Arena task success at 25% each. Some τ-Voice rows use fewer than three trials, and live Arena Elo can differ from the frozen value used in the index. The source-owned voice result does not enter weighted text-model scores; audio prices here are source observations, not canonical provider pricing.