Skip to main content
Live source snapshot

VoiceBench

A spoken question-answering suite spanning open-ended, knowledge, reasoning, instruction-following, and safety tasks.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
11 dataset subsets; human and synthesized speech
Primary metric
Benchmark-defined overall score
Owner
VoiceBench authors
Available evidence
Machine-readable public leaderboard

What it measures

Spoken reasoning
Does the system understand and answer spoken requests?

Available results

Protocol details
38 models
ModelArchitecture
NVIDIA Nemotron 3 Nano Omni 30B A3B
#1 · open
Vision + audio89.394.754.574.5882.3088.66
Ultravox-GLM-4P7
#2 · open
Audio-LLM88.864.874.304.5583.8276.26
Ultravox-GLM-4P7 (thinking)
#3 · open
Audio-LLM88.794.694.054.2187.1580.99
Whisper-v3-large + GPT-4o
#4 · closed
Cascaded87.804.804.474.6281.6976.51
Ultravox-GLM-4P6
#5 · open
Audio-LLM87.054.934.424.5780.4075.59
GPT-4o-Audio
#6 · closed
omni_hd86.754.784.494.5880.2576.02
GPT-4o-mini-Audio
#7 · closed
omni_hd82.844.754.244.4072.9072.90
Ultravox-v0.6-LLaMA-3.3-70B
#8 · open
Audio-LLM81.814.694.264.3869.2061.50
LFG-1
#9 · open
Audio-LLM79.634.603.714.0175.6074.85
Parakeet-TDT-0.6b-V2 + Qwen3-8B
#10 · open
Cascaded79.234.684.464.3559.1078.99
Whisper-v3-large + LLaMA-3.1-8B
#11 · open
Cascaded77.484.534.044.1662.4369.53
Kimi-Audio
#12 · open
omni_hd76.914.463.974.2062.1761.10
Whisper-v3-turbo + LLaMA-3.1-8B
#13 · open
Cascaded76.094.554.024.1262.0471.12
Ultravox-v0.5-LLaMA-3.1-8B
#14 · open
Audio-LLM74.864.594.114.2854.1666.51
Ultravox-v0.4.1-LLaMA-3.1-8B
#15 · open
Audio-LLM72.094.553.904.1247.1766.88
Baichuan-Omni-1.5
#16 · open
omni_hd71.324.504.054.0657.2554.54
MiniCPM-o
#17 · open
omni_hd71.234.424.153.9454.7849.25
Whisper-v3-turbo + LLaMA-3.2-3B
#18 · open
Cascaded71.024.453.824.0451.3769.71
Baichuan-Audio
#19 · open
omni_hd69.274.414.083.9253.1950.31
MERaLiON
#20 · open
Audio-LLM65.044.503.774.1234.9562.93
VITA-1.5
#21 · open
omni_hd64.534.213.663.4852.1538.14
Phi-4-multimodal
#22 · open
Vision + audio64.323.813.823.5642.1945.35
Ola
#23 · open
omni_hd59.424.122.973.1945.9739.57
Lyra-Base
#24 · open
omni_hd59.003.853.503.4249.7436.28
Ultravox-v0.5-LLaMA-3.2-1B
#26 · open
Audio-LLM57.464.043.573.4730.0345.56
DiVA
#27 · open
Audio-LLM57.393.673.543.7425.7639.15
GLM-4-Voice
#28 · open
omni_hd56.483.973.423.1839.7525.92
Qwen2-Audio
#29 · open
Audio-LLM55.803.743.433.0135.7226.33
Step-Audio
#31 · open
omni_hd50.844.133.092.9328.3327.96
Megrez-3B-Omni
#32 · open
omni_hd46.763.502.952.3427.0325.71
Ichigo
#33 · open
omni_hd45.573.793.172.8325.6321.59
Lyra-Mini
#34 · open
omni_hd45.262.992.692.5831.4220.91
Mair-hub-0.5B-Omni
#35 · open
omni_hd44.593.062.872.4825.6014.85
LLaMA-Omni
#36 · open
omni_hd41.123.703.462.9225.9314.87
VITA-1.0
#37 · open
omni_hd36.433.382.151.8725.7022.82
SLAM-Omni
#38 · open
omni_hd35.301.901.791.6026.0613.38
Mini-Omni2
#39 · open
omni_hd33.492.322.181.7924.2711.56
Mini-Omni
#40 · open
omni_hd30.421.952.021.6124.6913.58

Interpretation limit

Open-ended answers use an automatic model judge. Its overall score combines heterogeneous task scales using the benchmark owner’s method.