Skip to main content
BenchLM
Data

Massive Multi-discipline Multimodal Understanding (MMMU)

We show this table for reference; we do not rank on it.

Data verified 38 confirmed releases in the last 30 daysFollow model changes

A broad multimodal reasoning benchmark spanning charts, diagrams, tables, and academic visual question answering.

Benchmark score on MMMU — October 7, 2026

We compile the MMMU rows from provider self-reports and secondary reports. Qwen3.6 Plus leads the table at 86.0%, followed by Qwen3.5-122B-A10B (83.9%) and Qwen3.6-27B (82.9%). We do not use these results to rank models overall.

11 modelsMultimodal & GroundedRefreshingDisplay onlyUpdated October 7, 2026

Benchmark score results for MMMU
RankModel / configurationScoreParameters (B)Open / closed
1Qwen3.6 PlusAlibaba
86.0%
Not reportedClosed
2Qwen3.5-122B-A10BAlibaba
83.9%
Not reportedOpen
3Qwen3.6-27BAlibaba
82.9%
Not reportedOpen
4Qwen3.5-27BAlibaba
82.3%
Not reportedOpen
5Qwen3.6-35B-A3BAlibaba
81.7%
Not reportedOpen
6Qwen3.5-35B-A3BAlibaba
81.4%
Not reportedOpen
7Command A+Cohere
75.1%
Not reportedOpen
8Nemotron 3 Nano Omni 30B A3BNVIDIA
70.8%
Not reportedOpen
9LFM2.5-VL-3BLiquidAI
48.4%
Not reportedOpen
10ZAYA1-VL-8BZyphra
46.0%
Not reportedOpen
11LFM2.5-VL-450MLiquidAI
32.7%
Not reportedOpen

MMMU

BenchLM mirrors this Vals AI leaderboard for reference only. It is excluded from BenchLM overall and category rankings.

MMMU; source configurations remain separate from the canonical benchmark table.
Source rankSource model or configurationMMMU scoreSource setupParameters (B)Open / closed
1Claude Fable 5.1Anthropic90.64%Vals AIOriginal resultNot reportedClosed
2Claude Opus 5Anthropic89.88%Vals AIOriginal resultNot reportedClosed
3Claude Fable 5Anthropic89.31%Vals AIOriginal resultNot reportedClosed
4Gemini 3.8 FlashGoogle89.08%high reasoningOriginal resultNot reportedClosed
5Gemini 3.7 FlashGoogle88.96%high reasoningOriginal resultNot reportedClosed
6GPT-5.6 SolOpenAI88.84%max reasoningOriginal resultNot reportedClosed
7Gemini 3.6 FlashGoogle88.38%high reasoningOriginal resultNot reportedClosed
8Gemini 3.5 FlashGoogle88.27%high reasoningOriginal resultNot reportedClosed
9GPT-5.5OpenAI88.27%xhigh reasoningOriginal resultNot reportedClosed
10Gemini 3.1 Pro PreviewGoogle88.21%high reasoningOriginal resultNot reportedUnknown
11Kimi K3Moonshot AI88.15%Vals AIOriginal resultNot reportedPending
12Qwen3.8 MaxAlibaba88.03%Vals AIOriginal resultNot reportedOpen
13Gemini 3 Flash PreviewGoogle87.63%high reasoningOriginal resultNot reportedUnknown
14Gemini 3 Pro PreviewGoogle87.51%high reasoningOriginal resultNot reportedUnknown
15GPT-5.4OpenAI87.51%xhigh reasoningOriginal resultNot reportedClosed
16Muse SparkMeta87.40%Vals AIOriginal resultNot reportedClosed
17GPT-5.2OpenAI86.67%xhigh reasoningOriginal resultNot reportedClosed
18Muse Spark 1.1Meta86.59%xhigh reasoningOriginal resultNot reportedClosed
19Claude Opus 4.8Anthropic86.59%Vals AIOriginal resultNot reportedClosed
20GPT-5.6 TerraOpenAI86.47%max reasoningOriginal resultNot reportedClosed
21Kimi K2.6Moonshot AI86.30%Vals AIOriginal resultNot reportedOpen
22Muse Spark 1.2Meta86.13%xhigh reasoningOriginal resultNot reportedClosed
23GLM-5.3-FlashZ.AI86.01%max reasoningOriginal resultNot reportedOpen
24Claude Opus 4.7Anthropic85.55%Vals AIOriginal resultNot reportedClosed
25GPT-5.6 LunaOpenAI85.03%max reasoningOriginal resultNot reportedClosed
26Kimi K2.5 ThinkingMoonshot AI84.33%Vals AIOriginal resultNot reportedUnknown
27Qwen3.6 PlusAlibaba84.16%Vals AIOriginal resultNot reportedClosed
28Qwen3.8-27BAlibaba83.93%xhigh reasoningOriginal resultNot reportedOpen
29Claude Opus 4.6 (Adaptive)Anthropic83.87%Vals AIOriginal resultNot reportedClosed
30Gemini 3.5 Flash-LiteGoogle83.58%high reasoningOriginal resultNot reportedClosed
31Claude Sonnet 4.6Anthropic83.58%Vals AIOriginal resultNot reportedClosed
32Grok 4.20 0309 ReasoningxAI83.47%Vals AIOriginal resultNot reportedUnknown
33GPT-5.1OpenAI83.18%high reasoningOriginal resultNot reportedClosed
34Grok 4.3xAI83.06%Vals AIOriginal resultNot reportedClosed
35Claude Sonnet 5Anthropic83.01%Vals AIOriginal resultNot reportedClosed
36Claude Opus 4.5 ThinkingAnthropic82.95%Vals AIOriginal resultNot reportedClosed
37Gemini 3.1 Flash Lite PreviewGoogle82.49%high reasoningOriginal resultNot reportedUnknown
38Qwen3.5 FlashAlibaba81.91%Vals AIOriginal resultNot reportedClosed
39GPT-5OpenAI81.50%high reasoningOriginal resultNot reportedUnknown
40Gemini 2.5 Pro Exp 03 25Google81.34%Vals AIOriginal resultNot reportedUnknown
41MiniMax M3MiniMax81.16%Vals AIOriginal resultNot reportedOpen
42Claude Opus 4.5Anthropic81.10%Vals AIOriginal resultNot reportedClosed
43Gemini 2.5 Flash Preview 09 2025 ThinkingGoogle80.75%Vals AIOriginal resultNot reportedUnknown
44o3OpenAI80.42%high reasoningOriginal resultNot reportedClosed
45MiMo-V2.5Xiaomi80.00%Vals AIOriginal resultNot reportedClosed
46O4 MiniOpenAI79.67%high reasoningOriginal resultNot reportedUnknown
47Gemini 2.5 Flash Preview 09 2025Google79.48%Vals AIOriginal resultNot reportedUnknown
48Claude Sonnet 4.5 ThinkingAnthropic79.31%Vals AIOriginal resultNot reportedClosed
49GPT-5.4 miniOpenAI79.25%xhigh reasoningOriginal resultNot reportedClosed
50GPT-5 miniOpenAI78.91%high reasoningOriginal resultNot reportedClosed
51InklingThinking Machines Lab78.50%0.99 reasoningOriginal resultNot reportedOpen
52Claude Opus 4.1 20250805 ThinkingAnthropic77.51%Vals AIOriginal resultNot reportedUnknown
53o1OpenAI77.41%high reasoningOriginal resultNot reportedClosed
54Grok 4 0709xAI76.27%Vals AIOriginal resultNot reportedUnknown
55Gemini 2.5 Flash Lite Preview 09 2025 ThinkingGoogle75.43%Vals AIOriginal resultNot reportedUnknown
56Claude 3.7 Sonnet 20250219 ThinkingAnthropic75.10%Vals AIOriginal resultNot reportedUnknown
57Claude Sonnet 4 20250514 ThinkingAnthropic74.93%Vals AIOriginal resultNot reportedUnknown
58Claude Opus 4.1Anthropic73.72%Vals AIOriginal resultNot reportedUnknown
59GPT-5.4 nanoOpenAI73.58%high reasoningOriginal resultNot reportedClosed
60Claude Opus 4Anthropic73.31%Vals AIOriginal resultNot reportedUnknown
61Grok 4 Fast (Reasoning)xAI72.78%Vals AIOriginal resultNot reportedClosed
62Grok 4.1 Fast (Reasoning)xAI72.66%Vals AIOriginal resultNot reportedClosed
63Gemini 2.5 Flash Lite Preview 09 2025Google72.54%Vals AIOriginal resultNot reportedUnknown
64Claude Sonnet 4Anthropic72.39%Vals AIOriginal resultNot reportedUnknown
65GPT-4.1OpenAI72.39%high reasoningOriginal resultNot reportedClosed
66Gemini 2.5 Flash Preview 04 17 ThinkingGoogle71.92%Vals AIOriginal resultNot reportedUnknown
67Llama4 Maverick Instruct BasicFireworks AI71.69%Vals AIOriginal resultNot reportedUnknown
68Claude 3.7 SonnetAnthropic71.52%Vals AIOriginal resultNot reportedUnknown
69GPT-5 nanoOpenAI70.94%high reasoningOriginal resultNot reportedClosed
70GPT-4.1 miniOpenAI70.54%high reasoningOriginal resultNot reportedClosed
71Gemini 2.0 Flash 001Google69.79%Vals AIOriginal resultNot reportedUnknown
72Claude 3.5 SonnetAnthropic68.80%Vals AIOriginal resultNot reportedClosed
73Gemini 2.5 Flash Preview 04 17Google68.00%Vals AIOriginal resultNot reportedUnknown
74Mistral Large 2512Mistral AI66.19%Vals AIOriginal resultNot reportedUnknown
75Gemini 1.5 Pro 002Google65.51%Vals AIOriginal resultNot reportedUnknown
76Magistral Small 2509Mistral AI65.20%Vals AIOriginal resultNot reportedUnknown
77Magistral Medium 2509Mistral AI64.57%Vals AIOriginal resultNot reportedUnknown
78GPT-4oOpenAI64.01%high reasoningOriginal resultNot reportedClosed
79Grok 4.1 Fast Non ReasoningxAI63.70%Vals AIOriginal resultNot reportedUnknown
80Grok 4 Fast Non ReasoningxAI63.41%Vals AIOriginal resultNot reportedUnknown
81Mistral Medium 2505Mistral AI62.97%Vals AIOriginal resultNot reportedUnknown
82GPT-4oOpenAI62.16%high reasoningOriginal resultNot reportedClosed
83Grok 4.5xAI61.79%high reasoningOriginal resultNot reportedClosed
84Mistral Small 2503Mistral AI60.08%Vals AIOriginal resultNot reportedUnknown
85Meta Llama Llama 4 Scout 17B 16E InstructTogether AI58.75%Vals AIOriginal resultNot reportedUnknown
86Grok 2 Vision 1212xAI57.25%Vals AIOriginal resultNot reportedUnknown
87Gemini 1.5 Flash 002Google57.19%Vals AIOriginal resultNot reportedUnknown
88GPT-4o miniOpenAI56.56%high reasoningOriginal resultNot reportedClosed
89GPT-4.1 nanoOpenAI55.05%high reasoningOriginal resultNot reportedClosed
90Meta Llama Llama 3.2 90B Vision Instruct TurboTogether AI48.06%Vals AIOriginal resultNot reportedUnknown
91Claude Haiku 4.5 ThinkingAnthropic46.07%Vals AIOriginal resultNot reportedClosed
92Meta Llama Llama 3.2 11B Vision Instruct TurboTogether AI38.82%Vals AIOriginal resultNot reportedUnknown
93Qwen3.5 Plus ThinkingAlibaba22.77%Vals AIOriginal resultNot reportedUnknown

Among the reported MMMU rows, Qwen3.6 Plus is first at 86.0%. The third row is 3.1 points behind. The broader top-10 range is 40.0 points, so the table still separates the published systems.

11 models have been evaluated on MMMU. The benchmark falls in the Multimodal & Grounded category. MMMU is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About MMMU

Year

2024

Tasks

Multimodal academic reasoning

Format

Image + text question answering

Difficulty

Frontier multimodal

MMMU is the base benchmark family behind later MMMU-Pro variants. It measures whether a model can answer expert-style questions that require combining visual understanding with domain knowledge and reasoning.

Freshness and provenance

Version

MMMU 2024

Refresh cadence

Annual

Staleness state

Refreshing

Question availability

Public benchmark set

RefreshingDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does MMMU measure?

A broad multimodal reasoning benchmark spanning charts, diagrams, tables, and academic visual question answering.

Which model scores highest on MMMU?

Qwen3.6 Plus by Alibaba currently leads with a score of 86.0% on MMMU.

How many models are evaluated on MMMU?

11 AI models have published results on MMMU in the BenchLM catalog.

Last updated: October 7, 2026 · BenchLM version MMMU 2024

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 5,500+ readers.

One email each week. Unsubscribe anytime.