# VoiceBench: Voice Benchmark Profile

> A spoken question-answering suite spanning open-ended, knowledge, reasoning, instruction-following, and safety tasks.

- Scope: 11 dataset subsets; human and synthesized speech
- Measurement lanes: Spoken reasoning
- Primary metric: Benchmark-defined overall score
- Available evidence: Machine-readable public leaderboard
- Owner: VoiceBench authors
- Source snapshots refreshed: 2026-09-18

## Interpretation limit

Open-ended answers use an automatic model judge. Its overall score combines heterogeneous task scales using the benchmark owner’s method.

## Primary sources

- [Paper](https://arxiv.org/abs/2410.17196)
- [Code](https://github.com/MatthewCYM/VoiceBench)
- [Dataset](https://huggingface.co/datasets/VoiceBench/VoiceBench)
- [Owner page](https://matthewcym.github.io/VoiceBench/)

Canonical page: https://benchlm.ai/voice-benchmarks/voicebench
