# Big Bench Audio: Voice Benchmark Profile

> One thousand spoken reasoning questions test whether native audio models can answer correctly from audio and how quickly they begin replying.

- Scope: 1,000 English audio questions; four Big Bench Hard-derived categories with 250 questions each; 23 synthetic voices
- Measurement lanes: Spoken reasoning, Latency
- Primary metric: Speech reasoning accuracy (higher is better); time to first audio and measured cost are separate
- Available evidence: Published source table and public audio dataset
- Owner: Artificial Analysis
- Source snapshots refreshed: 2026-09-18

## Interpretation limit

The audio is synthetic and English-only. A model judge marks spoken answers correct or incorrect. The published summary table rounds reasoning accuracy to whole percentages; the cost-per-hour figure comes from a fixed 40-question subset and is not a provider price.

## Primary sources

- [Dataset](https://huggingface.co/datasets/ArtificialAnalysis/big_bench_audio)
- [Owner page](https://artificialanalysis.ai/speech-to-speech)

Canonical page: https://benchlm.ai/voice-benchmarks/big-bench-audio
