Big Bench Audio
One thousand spoken reasoning questions test whether native audio models can answer correctly from audio and how quickly they begin replying.
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Benchmark profile
- Scope
- 1,000 English audio questions; four Big Bench Hard-derived categories with 250 questions each; 23 synthetic voices
- Primary metric
- Speech reasoning accuracy (higher is better); time to first audio and measured cost are separate
- Owner
- Artificial Analysis
- Available evidence
- Published source table and public audio dataset
What it measures
How the protocol works
- 01
Questions cover formal fallacies, navigation, object counting, and web-of-lies reasoning.
- 02
Each question is delivered as audio; the tested model responds in audio.
- 03
A model judge compares the response with the official answer and question context.
- 04
Time to first audio and measured cost use the source’s separate API checks.
Available results
| Model | Reasoning (rounded) | First audio | Measured cost |
|---|---|---|---|
| StepAudio 3 Realtime · StepFun | ~100% | 8.83s | — |
| Qwen Audio 3.0 Realtime Plus · Alibaba Cloud | ~99% | 1.54s | $4.42/h |
| Qwen3.5 Omni Plus Realtime · Alibaba Cloud | ~99% | 2.64s | $0.00/h |
| Gemini 3.8 Live Extended Thinking (High) · Google | ~98% | 1.35s | $3.50/h |
| Step-Audio R1.1 (Realtime) · StepFun | ~98% | 1.53s | $0.00/h |
| Gemini 3.1 Flash Live High · Google | ~97% | 2.99s | $1.75/h |
| GPT-Realtime-2 (High) · OpenAI | ~97% | 1.14s | $4.14/h |
| Grok Voice Think Fast 1.0 · SpaceXAI | ~97% | 1.25s | $3.00/h |
| Grok Voice Think Fast 2.0 High · SpaceXAI | ~97% | 0.70s | $4.80/h |
| GPT-Realtime-2.1 High · OpenAI | ~96% | 1.21s | $10.75/h |
| Qwen Audio 3.0 Realtime Flash · Alibaba Cloud | ~96% | 1.55s | $4.77/h |
| GPT-Realtime-2 (Medium) · OpenAI | ~93% | 1.22s | $3.97/h |
| Grok Voice Fast 1.0 · SpaceXAI | ~93% | 0.78s | $3.00/h |
| Gemini 3.8 Live · Google | ~92% | 1.18s | $0.84/h |
| Gemini 2.5 Flash Native Audio Dialog Thinking · Google | ~91% | 3.87s | — |
| GPT-Live-1 (Astra, medium) · OpenAI | ~90% | 1.34s | $5.83/h |
| GPT-Live-1 (Sol, low) · OpenAI | ~89% | 1.24s | $4.47/h |
| Nova 2.0 Sonic (Mar 2026) · Amazon Bedrock | ~88% | 1.14s | — |
| GPT-Realtime-2.1 Minimal · OpenAI | ~87% | 0.97s | $11.31/h |
| Deepslate Opal · Deepslate | ~85% | 0.44s | $6.48/h |
| GPT Realtime (Aug '25) · OpenAI | ~83% | 0.98s | $11.08/h |
| GPT-Realtime-1.5 · OpenAI | ~81% | 0.81s | $11.44/h |
| GPT-Realtime-2.1 Mini High · OpenAI | ~75% | 4.28s | $3.45/h |
| GPT-Realtime-2 (Minimal) · OpenAI | ~72% | 1.12s | $3.07/h |
| Gemini 3.1 Flash Live Minimal · Google | ~71% | 0.96s | $1.50/h |
| Gemini 2.5 Flash Native Audio Dialog · Google | ~69% | 0.63s | $1.42/h |
| GPT-4o mini Realtime (Dec 2024) · OpenAI | ~69% | 1.27s | $5.75/h |
| Higgs Realtime · Boson AI | ~69% | 1.47s | — |
| GPT Realtime Mini (Oct '25) · OpenAI | ~64% | 0.81s | $3.04/h |
| GPT-Realtime-2.1 Mini Minimal · OpenAI | ~63% | 0.85s | $4.60/h |
| Qwen3 Omni Flash · Alibaba Cloud | ~59% | 4.82s | $1.77/h |
| Qwen3.5 Omni Flash Realtime · Alibaba Cloud | ~59% | 0.79s | $0.00/h |
| Qwen3 Omni Realtime · Alibaba Cloud | ~57% | 0.88s | $2.26/h |
| Raon SpeechChat · Krafton | ~57% | 0.04s | — |
| GPT-4o audio chatcompletions · OpenAI | ~54% | 3.38s | $0.00/h |
| Freeze-Omni · VITA | ~33% | — | — |
| Nemotron Voicechat · NVIDIA | ~27% | — | — |
| PersonaPlex · NVIDIA | ~19% | — | — |
| FLM-Audio · Cofe AI | ~16% | — | — |
| Moshi · Kyutai | ~4% | — | — |