Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes

Massive Multi-discipline Multimodal Understanding Pro (MMMU-Pro)

A harder multimodal benchmark for frontier models that combines text with images, diagrams, charts, and academic visual reasoning tasks.

Data verified 33 confirmed releases in the last 30 daysSee provider release alerts

Top models on MMMU-Pro — September 15, 2026

As of September 15, 2026, Gemini 3.1 Pro leads the MMMU-Pro leaderboard with 83.9% , followed by Gemini 3.5 Flash (83.6%) and GPT-5.6 Sol (83%).

40 modelsMultimodal & Grounded40% of category scoreRefreshingUpdated September 15, 2026

Leaderboard (40 models)

Score
1
Gemini 3.1 ProGoogle · Closed
83.9%
2
Gemini 3.5 FlashGoogle · Closed
83.6%
3
GPT-5.6 SolOpenAI · Closed
83%
4
Qwen3.8 MaxAlibaba · Open weight
82.3%
5
Kimi K3Moonshot AI · Closed
81.6%
6
Seed 2.1 ProByteDance · Closed
81.6%
7
GPT-5.5OpenAI · Closed
81.2%
8
GPT-5.4OpenAI · Closed
81.2%
9
Gemini 3 ProGoogle · Closed
81%
10
GPT-5.6 TerraOpenAI · Closed
80.7%
11
Muse SparkMeta · Closed
80.4%
12
Seed 2.1 TurboByteDance · Closed
80.1%
13
GPT-5.2OpenAI · Closed
79.5%
14
Kimi K2.6Moonshot AI · Open weight
79.4%
15
dots3-note PreviewDots Studio · Open weight
79.1%
16
Qwen3.7 PlusAlibaba · Closed
79%
17
Qwen3.5 397BAlibaba · Open weight
79%
18
Qwen3.6 PlusAlibaba · Closed
78.8%
19
Kimi K2.5Moonshot AI · Open weight
78.5%
20
Kimi K2.5 (Reasoning)Moonshot AI · Closed
78.5%
21
GPT-5.6 LunaOpenAI · Closed
78.4%
22
MiniMax M3MiniMax · Open weight
78.1%
23
Grok 4.3xAI · Closed
78.1%
24
MiMo-V2.5Xiaomi · Closed
77.9%
25
Claude Opus 4.6Anthropic · Closed
77.3%
26
Gemma 4 31BGoogle · Open weight
76.9%
27
GPT-5.4 miniOpenAI · Closed
76.6%
28
Qwen3.6-27BAlibaba · Open weight
75.8%
29
Qwen3.6-35B-A3BAlibaba · Open weight
75.3%
30
Grok 4.20xAI · Closed
75.2%
31
Inkling-SmallThinking Machines Lab · Open weight
74%
32
Muse Glimmer 30BMeta · Open weight
74%
33
Gemma 4 26B A4BGoogle · Open weight
73.8%
34
InklingThinking Machines Lab · Open weight
73.5%
35
Interfaze BetaInterfaze · Closed
71.1%
36
Claude Opus 4.5Anthropic · Closed
70.6%
37
Gemma 4 12BGoogle · Open weight
69.1%
38
GPT-5.4 nanoOpenAI · Closed
66.1%
39
Command A+Cohere · Open weight
63%
40
LFM2.5-VL-3BLiquidAI · Open weight
30.5%

According to BenchLM.ai, Gemini 3.1 Pro leads the MMMU-Pro benchmark with a score of 83.9%, followed by Gemini 3.5 Flash (83.6%) and GPT-5.6 Sol (83%). The top models are clustered within 0.9 points, suggesting this benchmark is nearing saturation for frontier models.

40 models have been evaluated on MMMU-Pro. The benchmark falls in the Multimodal & Grounded category. This category carries a 12% weight in BenchLM.ai's overall scoring system. Within that category, MMMU-Pro contributes 40% of the category score, so strong performance here directly affects a model's overall ranking.

About MMMU-Pro

Year

2024

Tasks

Multimodal academic reasoning

Format

Image + text question answering

Difficulty

Frontier multimodal

MMMU-Pro extends the original MMMU setup with more difficult multimodal questions and stronger separation at the top end of the model market.

BenchLM freshness & provenance

Version

MMMU-Pro 2024

Refresh cadence

Annual

Staleness state

Refreshing

Question availability

Public benchmark set

Refreshing

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does MMMU-Pro measure?

A harder multimodal benchmark for frontier models that combines text with images, diagrams, charts, and academic visual reasoning tasks.

Which model scores highest on MMMU-Pro?

Gemini 3.1 Pro by Google currently leads with a score of 83.9% on MMMU-Pro.

How many models are evaluated on MMMU-Pro?

40 AI models have been evaluated on MMMU-Pro on BenchLM.

Last updated: September 15, 2026 · BenchLM version MMMU-Pro 2024

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.