Live source snapshot
MMAU-Pro
Tests harder audio reasoning across speech, sound, music, spatial audio, multiple clips, voice chat, and instruction following.
Source snapshots refreshed
Five measurement lanesVOICE / S2S
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Benchmark profile
- Scope
- 5,305 human-curated questions; 49 skills; audio up to 10 minutes
- Primary metric
- Average accuracy plus task-family accuracy
- Owner
- MMAU-Pro authors
- Available evidence
- Machine-readable benchmark-owner leaderboard
What it measures
Spoken reasoning
Does the system understand and answer spoken requests?
Conversation dynamics
Can it handle turns, interruptions, ambiguity, and state?
Available results
| Rank | Model | Average | Voice | Speech | Sound | Music |
|---|---|---|---|---|---|---|
| 1 | Gemini-3.1 Pro Preview | 72.2% | 82.1% | 88.9% | 60.0% | 76.0% |
| 2 | Gemini-3.7 Flash | 70.4% | 87.9% | 89.7% | 56.4% | 71.7% |
| 3 | Gemini-3.6 Flash | 69.6% | 84.8% | 86.6% | 55.3% | 71.1% |
| 4 | Gemini-3.5 Flash | 68.5% | 73.8% | 85.0% | 56.8% | 70.7% |
| 5 | Qwen3-Omni-30B-A3B-Instruct | 62.9% | 66.2% | 73.2% | 47.4% | 74.9% |
| 6 | Gemini-2.5 Flash | 59.2% | 71.7% | 73.4% | 51.9% | 64.9% |
| 7 | gpt-audio | 58.8% | 49.3% | 73.1% | 46.2% | 67.3% |
| 8 | Gemini-2.0 Flash | 55.7% | 68.6% | 69.5% | 48.4% | 56.9% |
| 9 | GPT4o-Audio | 52.5% | 57.5% | 68.2% | 44.7% | 63.1% |
| 10 | Qwen2.5-Omni-7B | 52.2% | 60.0% | 57.4% | 47.6% | 61.5% |
| 11 | Audio Flamingo 3 | 51.7% | 58.6% | 58.8% | 55.9% | 61.7% |
| 12 | GPT4o-mini-Audio | 48.3% | 52.7% | 66.1% | 40.2% | 59.7% |
| 13 | Kimi-Audio-7B-Instruct | 48.3% | 32.1% | 55.8% | 45.9% | 50.4% |
| 14 | Ming-Lite-Omni-1.5 | 47.4% | 44.5% | 49.1% | 47.9% | 56.2% |
| 15 | Kimi-Audio | 46.6% | 50.6% | 52.2% | 46.0% | 57.6% |
| 16 | Qwen2.5-Omni-3B | 46.1% | 46.5% | 53.9% | 38.5% | 60.3% |
| 17 | Audio Flamingo 2 | 42.6% | 37.2% | 43.0% | 39.5% | 55.7% |
| 18 | DeSTA2.5-Audio | 40.6% | 51.0% | 49.9% | 35.7% | 48.2% |
| 19 | Gemma-3n-E4B-it | 39.7% | 58.3% | 44.9% | 42.4% | 46.4% |
| 20 | SALMONN 13B | 39.6% | 53.2% | 37.3% | 43.6% | 47.2% |
| 21 | Audio-Reasoner | 39.5% | 43.4% | 44.0% | 34.2% | 50.1% |
| 22 | Phi4-MM | 38.7% | 42.7% | 47.6% | 25.7% | 47.8% |
| 23 | DeSTA2 | 36.7% | 54.8% | 46.5% | 31.0% | 43.3% |
| 24 | Gemma-3n-E2B-it | 35.4% | 51.4% | 41.3% | 40.1% | 44.1% |
| 25 | Caption + GPT4o | 35.3% | 38.6% | 38.4% | 38.6% | 40.6% |
| 26 | SALMONN 7B | 34.5% | 36.5% | 38.3% | 32.2% | 44.9% |
| 27 | R1-AQA | 34.1% | 32.7% | 33.7% | 47.9% | 31.9% |
| 28 | Baichuan-Omni-1.5 | 33.9% | 40.0% | 36.5% | 34.6% | 32.5% |
| 29 | Captions + Qwen235B-A22B | 33.7% | 35.6% | 36.1% | 36.4% | 41.3% |
| 30 | GAMA | 33.2% | 28.4% | 29.8% | 45.4% | 41.2% |
| 31 | Mellow | 27.5% | 28.3% | 27.9% | 27.6% | 32.9% |
| 32 | BAT | 24.8% | 24.5% | 25.9% | 28.9% | 22.7% |
Interpretation limit
The overall average spans heterogeneous audio tasks. Inspect speech, voice, spatial, multi-audio, and instruction-following columns before using it to choose a system.