Skip to main content
Radar

Five or fewer confirmed AI changes, with original sources, on mornings when something changed.A free source-linked morning brief.

See the free Radar Brief
Live source snapshot

MMAU-Pro

Tests harder audio reasoning across speech, sound, music, spatial audio, multiple clips, voice chat, and instruction following.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
5,305 human-curated questions; 49 skills; audio up to 10 minutes
Primary metric
Average accuracy plus task-family accuracy
Owner
MMAU-Pro authors
Available evidence
Machine-readable benchmark-owner leaderboard

What it measures

Spoken reasoning
Does the system understand and answer spoken requests?
Conversation dynamics
Can it handle turns, interruptions, ambiguity, and state?

Available results

MMAU-Pro average and selected audio-task accuracy by model.
RankModelAverageVoiceSpeechSoundMusic
1Gemini-3.1 Pro Preview72.2%82.1%88.9%60.0%76.0%
2Gemini-3.7 Flash70.4%87.9%89.7%56.4%71.7%
3Gemini-3.6 Flash69.6%84.8%86.6%55.3%71.1%
4Gemini-3.5 Flash68.5%73.8%85.0%56.8%70.7%
5Qwen3-Omni-30B-A3B-Instruct62.9%66.2%73.2%47.4%74.9%
6Gemini-2.5 Flash59.2%71.7%73.4%51.9%64.9%
7gpt-audio58.8%49.3%73.1%46.2%67.3%
8Gemini-2.0 Flash55.7%68.6%69.5%48.4%56.9%
9GPT4o-Audio52.5%57.5%68.2%44.7%63.1%
10Qwen2.5-Omni-7B52.2%60.0%57.4%47.6%61.5%
11Audio Flamingo 351.7%58.6%58.8%55.9%61.7%
12GPT4o-mini-Audio48.3%52.7%66.1%40.2%59.7%
13Kimi-Audio-7B-Instruct48.3%32.1%55.8%45.9%50.4%
14Ming-Lite-Omni-1.547.4%44.5%49.1%47.9%56.2%
15Kimi-Audio46.6%50.6%52.2%46.0%57.6%
16Qwen2.5-Omni-3B46.1%46.5%53.9%38.5%60.3%
17Audio Flamingo 242.6%37.2%43.0%39.5%55.7%
18DeSTA2.5-Audio40.6%51.0%49.9%35.7%48.2%
19Gemma-3n-E4B-it39.7%58.3%44.9%42.4%46.4%
20SALMONN 13B39.6%53.2%37.3%43.6%47.2%
21Audio-Reasoner39.5%43.4%44.0%34.2%50.1%
22Phi4-MM38.7%42.7%47.6%25.7%47.8%
23DeSTA236.7%54.8%46.5%31.0%43.3%
24Gemma-3n-E2B-it35.4%51.4%41.3%40.1%44.1%
25Caption + GPT4o35.3%38.6%38.4%38.6%40.6%
26SALMONN 7B34.5%36.5%38.3%32.2%44.9%
27R1-AQA34.1%32.7%33.7%47.9%31.9%
28Baichuan-Omni-1.533.9%40.0%36.5%34.6%32.5%
29Captions + Qwen235B-A22B33.7%35.6%36.1%36.4%41.3%
30GAMA33.2%28.4%29.8%45.4%41.2%
31Mellow27.5%28.3%27.9%27.6%32.9%
32BAT24.8%24.5%25.9%28.9%22.7%

Interpretation limit

The overall average spans heterogeneous audio tasks. Inspect speech, voice, spatial, multi-audio, and instruction-following columns before using it to choose a system.