Live source snapshot
MMAU
Measures expert-level understanding and reasoning across speech, environmental sound, and music.
Source snapshots refreshed
Five measurement lanesVOICE / S2S
Reasoning, task completion, conversation dynamics, experience, and latency stay separate.
Benchmark profile
- Scope
- 10,000 audio clips; 27 tasks; speech, sound, and music domains
- Primary metric
- Test accuracy by domain and overall average
- Owner
- MMAU authors
- Available evidence
- Machine-readable benchmark-owner leaderboard
What it measures
Spoken reasoning
Does the system understand and answer spoken requests?
Available results
| Rank | Model | Average | Sound | Music | Speech |
|---|---|---|---|---|---|
| 1 | Audio-Thinker | 75.98% | 78.80% | 73.80% | 75.16% |
| 2 | Nova 2 Omni | 75.28% | 77.87% | 66.37% | 81.82% |
| 3 | Step-Audio-2 | 73.86% | 80.60% | 68.23% | 72.75% |
| 4 | MiMo-Audio | 72.59% | 77.20% | 69.73% | 70.77% |
| 5 | Audio Flamingo 3 | 72.42% | 75.83% | 74.47% | 66.97% |
| 6 | Qwen2.5-Omni | 71.00% | 76.77% | 67.33% | 68.90% |
| 7 | Step-Audio-2-mini | 70.23% | 75.57% | 66.85% | 66.49% |
| 8 | Gemini 2.5 Pro | 69.36% | 70.63% | 64.77% | 72.67% |
| 9 | Gemini 2.5 Flash | 67.39% | 69.50% | 69.40% | 68.27% |
| 10 | Gemini 2.0 Flash | 67.03% | 68.93% | 59.30% | 72.87% |
| 11 | DeSTA2.5-Audio | 65.21% | 66.83% | 57.10% | 71.94% |
| 12 | Kimi-Audio | 64.40% | 70.70% | 65.93% | 56.57% |
| 13 | Audio Reasoner | 63.78% | 67.27% | 61.53% | 62.53% |
| 14 | Phi-4-multimodal | 62.81% | 62.67% | 61.97% | 63.80% |
| 15 | Gemini 2.5 Flash Lite | 61.61% | 62.50% | 54.87% | 67.47% |
| 16 | Audio Flamingo 2 | 61.06% | 68.13% | 70.20% | 44.87% |
| 17 | GPT-4o Audio | 60.82% | 63.20% | 49.93% | 69.33% |
| 18 | Qwen2-Audio-Instruct | 57.40% | 61.17% | 55.67% | 55.37% |
| 19 | Gemma 3n | 55.20% | 50.27% | 53.20% | 62.13% |
| 20 | Gemma 3n | 52.06% | 47.47% | 51.63% | 57.07% |
| 21 | GPT-4o mini Audio | 51.03% | 49.67% | 35.97% | 67.47% |
| 22 | M2UGen | 39.76% | 44.97% | 38.53% | 35.77% |
| 23 | MusiLingo | 38.29% | 41.93% | 41.23% | 31.70% |
| 24 | SALMONN | 36.23% | 42.10% | 37.83% | 28.77% |
| 25 | MuLLaMa | 25.91% | 30.97% | 29.67% | 17.10% |
| 26 | GAMA-IT | 22.22% | 32.73% | 22.37% | 11.57% |
| 27 | GAMA | 21.68% | 30.73% | 17.33% | 16.97% |
| 28 | LTU | 17.23% | 20.67% | 15.68% | 15.33% |
| 29 | Audio Flamingo Chat | 15.59% | 23.33% | 15.77% | 7.67% |
Interpretation limit
The refreshed table uses the parsed MMAU-v05.15.25 test results. It does not merge older benchmark versions or unverified community-reported rows.