Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes
Live source snapshot

MMAU

Measures expert-level understanding and reasoning across speech, environmental sound, and music.

Source snapshots refreshed
Five measurement lanesVOICE / S2S

Reasoning, task completion, conversation dynamics, experience, and latency stay separate.

Back to directory

Benchmark profile

Scope
10,000 audio clips; 27 tasks; speech, sound, and music domains
Primary metric
Test accuracy by domain and overall average
Owner
MMAU authors
Available evidence
Machine-readable benchmark-owner leaderboard

What it measures

Spoken reasoning
Does the system understand and answer spoken requests?

Available results

MMAU-v05.15.25 test accuracy across sound, music, and speech.
RankModelAverageSoundMusicSpeech
1Audio-Thinker75.98%78.80%73.80%75.16%
2Nova 2 Omni75.28%77.87%66.37%81.82%
3Step-Audio-273.86%80.60%68.23%72.75%
4MiMo-Audio72.59%77.20%69.73%70.77%
5Audio Flamingo 372.42%75.83%74.47%66.97%
6Qwen2.5-Omni71.00%76.77%67.33%68.90%
7Step-Audio-2-mini70.23%75.57%66.85%66.49%
8Gemini 2.5 Pro69.36%70.63%64.77%72.67%
9Gemini 2.5 Flash67.39%69.50%69.40%68.27%
10Gemini 2.0 Flash67.03%68.93%59.30%72.87%
11DeSTA2.5-Audio65.21%66.83%57.10%71.94%
12Kimi-Audio64.40%70.70%65.93%56.57%
13Audio Reasoner63.78%67.27%61.53%62.53%
14Phi-4-multimodal62.81%62.67%61.97%63.80%
15Gemini 2.5 Flash Lite61.61%62.50%54.87%67.47%
16Audio Flamingo 261.06%68.13%70.20%44.87%
17GPT-4o Audio60.82%63.20%49.93%69.33%
18Qwen2-Audio-Instruct57.40%61.17%55.67%55.37%
19Gemma 3n55.20%50.27%53.20%62.13%
20Gemma 3n52.06%47.47%51.63%57.07%
21GPT-4o mini Audio51.03%49.67%35.97%67.47%
22M2UGen39.76%44.97%38.53%35.77%
23MusiLingo38.29%41.93%41.23%31.70%
24SALMONN36.23%42.10%37.83%28.77%
25MuLLaMa25.91%30.97%29.67%17.10%
26GAMA-IT22.22%32.73%22.37%11.57%
27GAMA21.68%30.73%17.33%16.97%
28LTU17.23%20.67%15.68%15.33%
29Audio Flamingo Chat15.59%23.33%15.77%7.67%

Interpretation limit

The refreshed table uses the parsed MMAU-v05.15.25 test results. It does not merge older benchmark versions or unverified community-reported rows.