Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes
Live source snapshot

Big Bench Audio

One thousand spoken reasoning questions test whether native audio models can answer correctly from audio and how quickly they begin replying.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
1,000 English audio questions; four Big Bench Hard-derived categories with 250 questions each; 23 synthetic voices
Primary metric
Speech reasoning accuracy (higher is better); time to first audio and measured cost are separate
Owner
Artificial Analysis
Available evidence
Published source table and public audio dataset

What it measures

Spoken reasoning
Does the system understand and answer spoken requests?
Latency
How long do responses, tool calls, and full tasks take?

How the protocol works

  1. 01

    Questions cover formal fallacies, navigation, object counting, and web-of-lies reasoning.

  2. 02

    Each question is delivered as audio; the tested model responds in audio.

  3. 03

    A model judge compares the response with the official answer and question context.

  4. 04

    Time to first audio and measured cost use the source’s separate API checks.

Available results

Big Bench Audio speech reasoning, time to first audio, and measured subset cost. Published table precision.
ModelReasoning (rounded)First audioMeasured cost
StepAudio 3 Realtime · StepFun~100%8.83s
Qwen Audio 3.0 Realtime Plus · Alibaba Cloud~99%1.54s$4.42/h
Qwen3.5 Omni Plus Realtime · Alibaba Cloud~99%2.64s$0.00/h
Gemini 3.8 Live Extended Thinking (High) · Google~98%1.35s$3.50/h
Step-Audio R1.1 (Realtime) · StepFun~98%1.53s$0.00/h
Gemini 3.1 Flash Live High · Google~97%2.99s$1.75/h
GPT-Realtime-2 (High) · OpenAI~97%1.14s$4.14/h
Grok Voice Think Fast 1.0 · SpaceXAI~97%1.25s$3.00/h
Grok Voice Think Fast 2.0 High · SpaceXAI~97%0.70s$4.80/h
GPT-Realtime-2.1 High · OpenAI~96%1.21s$10.75/h
Qwen Audio 3.0 Realtime Flash · Alibaba Cloud~96%1.55s$4.77/h
GPT-Realtime-2 (Medium) · OpenAI~93%1.22s$3.97/h
Grok Voice Fast 1.0 · SpaceXAI~93%0.78s$3.00/h
Gemini 3.8 Live · Google~92%1.18s$0.84/h
Gemini 2.5 Flash Native Audio Dialog Thinking · Google~91%3.87s
GPT-Live-1 (Astra, medium) · OpenAI~90%1.34s$5.83/h
GPT-Live-1 (Sol, low) · OpenAI~89%1.24s$4.47/h
Nova 2.0 Sonic (Mar 2026) · Amazon Bedrock~88%1.14s
GPT-Realtime-2.1 Minimal · OpenAI~87%0.97s$11.31/h
Deepslate Opal · Deepslate~85%0.44s$6.48/h
GPT Realtime (Aug '25) · OpenAI~83%0.98s$11.08/h
GPT-Realtime-1.5 · OpenAI~81%0.81s$11.44/h
GPT-Realtime-2.1 Mini High · OpenAI~75%4.28s$3.45/h
GPT-Realtime-2 (Minimal) · OpenAI~72%1.12s$3.07/h
Gemini 3.1 Flash Live Minimal · Google~71%0.96s$1.50/h
Gemini 2.5 Flash Native Audio Dialog · Google~69%0.63s$1.42/h
GPT-4o mini Realtime (Dec 2024) · OpenAI~69%1.27s$5.75/h
Higgs Realtime · Boson AI~69%1.47s
GPT Realtime Mini (Oct '25) · OpenAI~64%0.81s$3.04/h
GPT-Realtime-2.1 Mini Minimal · OpenAI~63%0.85s$4.60/h
Qwen3 Omni Flash · Alibaba Cloud~59%4.82s$1.77/h
Qwen3.5 Omni Flash Realtime · Alibaba Cloud~59%0.79s$0.00/h
Qwen3 Omni Realtime · Alibaba Cloud~57%0.88s$2.26/h
Raon SpeechChat · Krafton~57%0.04s
GPT-4o audio chatcompletions · OpenAI~54%3.38s$0.00/h
Freeze-Omni · VITA~33%
Nemotron Voicechat · NVIDIA~27%
PersonaPlex · NVIDIA~19%
FLM-Audio · Cofe AI~16%
Moshi · Kyutai~4%

Interpretation limit

The audio is synthetic and English-only. A model judge marks spoken answers correct or incorrect. The published summary table rounds reasoning accuracy to whole percentages; the cost-per-hour figure comes from a fixed 40-question subset and is not a provider price.