Massive Multi-discipline Multimodal Understanding (MMMU)
We show this table for reference; we do not rank on it.
A broad multimodal reasoning benchmark spanning charts, diagrams, tables, and academic visual question answering.
Benchmark score on MMMU — October 7, 2026
We compile the MMMU rows from provider self-reports and secondary reports. Qwen3.6 Plus leads the table at 86.0%, followed by Qwen3.5-122B-A10B (83.9%) and Qwen3.6-27B (82.9%). We do not use these results to rank models overall.
Qwen3.6 Plus
Alibaba
Qwen3.5-122B-A10B
Alibaba
Qwen3.6-27B
Alibaba
11 modelsMultimodal & GroundedRefreshingDisplay onlyUpdated October 7, 2026
| Rank | Model / configuration | Score | Parameters (B) | Open / closed |
|---|---|---|---|---|
| 1 | Qwen3.6 PlusAlibaba | 86.0% | Not reported | Closed |
| 2 | Qwen3.5-122B-A10BAlibaba | 83.9% | Not reported | Open |
| 3 | Qwen3.6-27BAlibaba | 82.9% | Not reported | Open |
| 4 | Qwen3.5-27BAlibaba | 82.3% | Not reported | Open |
| 5 | Qwen3.6-35B-A3BAlibaba | 81.7% | Not reported | Open |
| 6 | Qwen3.5-35B-A3BAlibaba | 81.4% | Not reported | Open |
| 7 | Command A+Cohere | 75.1% | Not reported | Open |
| 8 | Nemotron 3 Nano Omni 30B A3BNVIDIA | 70.8% | Not reported | Open |
| 9 | LFM2.5-VL-3BLiquidAI | 48.4% | Not reported | Open |
| 10 | ZAYA1-VL-8BZyphra | 46.0% | Not reported | Open |
| 11 | LFM2.5-VL-450MLiquidAI | 32.7% | Not reported | Open |
MMMU
BenchLM mirrors this Vals AI leaderboard for reference only. It is excluded from BenchLM overall and category rankings.
| Source rank | Source model or configuration | MMMU score | Source setup | Parameters (B) | Open / closed |
|---|---|---|---|---|---|
| 1 | Claude Fable 5.1Anthropic | 90.64% | Vals AIOriginal result | Not reported | Closed |
| 2 | Claude Opus 5Anthropic | 89.88% | Vals AIOriginal result | Not reported | Closed |
| 3 | Claude Fable 5Anthropic | 89.31% | Vals AIOriginal result | Not reported | Closed |
| 4 | Gemini 3.8 FlashGoogle | 89.08% | high reasoningOriginal result | Not reported | Closed |
| 5 | Gemini 3.7 FlashGoogle | 88.96% | high reasoningOriginal result | Not reported | Closed |
| 6 | GPT-5.6 SolOpenAI | 88.84% | max reasoningOriginal result | Not reported | Closed |
| 7 | Gemini 3.6 FlashGoogle | 88.38% | high reasoningOriginal result | Not reported | Closed |
| 8 | Gemini 3.5 FlashGoogle | 88.27% | high reasoningOriginal result | Not reported | Closed |
| 9 | GPT-5.5OpenAI | 88.27% | xhigh reasoningOriginal result | Not reported | Closed |
| 10 | Gemini 3.1 Pro PreviewGoogle | 88.21% | high reasoningOriginal result | Not reported | Unknown |
| 11 | Kimi K3Moonshot AI | 88.15% | Vals AIOriginal result | Not reported | Pending |
| 12 | Qwen3.8 MaxAlibaba | 88.03% | Vals AIOriginal result | Not reported | Open |
| 13 | Gemini 3 Flash PreviewGoogle | 87.63% | high reasoningOriginal result | Not reported | Unknown |
| 14 | Gemini 3 Pro PreviewGoogle | 87.51% | high reasoningOriginal result | Not reported | Unknown |
| 15 | GPT-5.4OpenAI | 87.51% | xhigh reasoningOriginal result | Not reported | Closed |
| 16 | Muse SparkMeta | 87.40% | Vals AIOriginal result | Not reported | Closed |
| 17 | GPT-5.2OpenAI | 86.67% | xhigh reasoningOriginal result | Not reported | Closed |
| 18 | Muse Spark 1.1Meta | 86.59% | xhigh reasoningOriginal result | Not reported | Closed |
| 19 | Claude Opus 4.8Anthropic | 86.59% | Vals AIOriginal result | Not reported | Closed |
| 20 | GPT-5.6 TerraOpenAI | 86.47% | max reasoningOriginal result | Not reported | Closed |
| 21 | Kimi K2.6Moonshot AI | 86.30% | Vals AIOriginal result | Not reported | Open |
| 22 | Muse Spark 1.2Meta | 86.13% | xhigh reasoningOriginal result | Not reported | Closed |
| 23 | GLM-5.3-FlashZ.AI | 86.01% | max reasoningOriginal result | Not reported | Open |
| 24 | Claude Opus 4.7Anthropic | 85.55% | Vals AIOriginal result | Not reported | Closed |
| 25 | GPT-5.6 LunaOpenAI | 85.03% | max reasoningOriginal result | Not reported | Closed |
| 26 | Kimi K2.5 ThinkingMoonshot AI | 84.33% | Vals AIOriginal result | Not reported | Unknown |
| 27 | Qwen3.6 PlusAlibaba | 84.16% | Vals AIOriginal result | Not reported | Closed |
| 28 | Qwen3.8-27BAlibaba | 83.93% | xhigh reasoningOriginal result | Not reported | Open |
| 29 | Claude Opus 4.6 (Adaptive)Anthropic | 83.87% | Vals AIOriginal result | Not reported | Closed |
| 30 | Gemini 3.5 Flash-LiteGoogle | 83.58% | high reasoningOriginal result | Not reported | Closed |
| 31 | Claude Sonnet 4.6Anthropic | 83.58% | Vals AIOriginal result | Not reported | Closed |
| 32 | Grok 4.20 0309 ReasoningxAI | 83.47% | Vals AIOriginal result | Not reported | Unknown |
| 33 | GPT-5.1OpenAI | 83.18% | high reasoningOriginal result | Not reported | Closed |
| 34 | Grok 4.3xAI | 83.06% | Vals AIOriginal result | Not reported | Closed |
| 35 | Claude Sonnet 5Anthropic | 83.01% | Vals AIOriginal result | Not reported | Closed |
| 36 | Claude Opus 4.5 ThinkingAnthropic | 82.95% | Vals AIOriginal result | Not reported | Closed |
| 37 | Gemini 3.1 Flash Lite PreviewGoogle | 82.49% | high reasoningOriginal result | Not reported | Unknown |
| 38 | Qwen3.5 FlashAlibaba | 81.91% | Vals AIOriginal result | Not reported | Closed |
| 39 | GPT-5OpenAI | 81.50% | high reasoningOriginal result | Not reported | Unknown |
| 40 | Gemini 2.5 Pro Exp 03 25Google | 81.34% | Vals AIOriginal result | Not reported | Unknown |
| 41 | MiniMax M3MiniMax | 81.16% | Vals AIOriginal result | Not reported | Open |
| 42 | Claude Opus 4.5Anthropic | 81.10% | Vals AIOriginal result | Not reported | Closed |
| 43 | Gemini 2.5 Flash Preview 09 2025 ThinkingGoogle | 80.75% | Vals AIOriginal result | Not reported | Unknown |
| 44 | o3OpenAI | 80.42% | high reasoningOriginal result | Not reported | Closed |
| 45 | MiMo-V2.5Xiaomi | 80.00% | Vals AIOriginal result | Not reported | Closed |
| 46 | O4 MiniOpenAI | 79.67% | high reasoningOriginal result | Not reported | Unknown |
| 47 | Gemini 2.5 Flash Preview 09 2025Google | 79.48% | Vals AIOriginal result | Not reported | Unknown |
| 48 | Claude Sonnet 4.5 ThinkingAnthropic | 79.31% | Vals AIOriginal result | Not reported | Closed |
| 49 | GPT-5.4 miniOpenAI | 79.25% | xhigh reasoningOriginal result | Not reported | Closed |
| 50 | GPT-5 miniOpenAI | 78.91% | high reasoningOriginal result | Not reported | Closed |
| 51 | InklingThinking Machines Lab | 78.50% | 0.99 reasoningOriginal result | Not reported | Open |
| 52 | Claude Opus 4.1 20250805 ThinkingAnthropic | 77.51% | Vals AIOriginal result | Not reported | Unknown |
| 53 | o1OpenAI | 77.41% | high reasoningOriginal result | Not reported | Closed |
| 54 | Grok 4 0709xAI | 76.27% | Vals AIOriginal result | Not reported | Unknown |
| 55 | Gemini 2.5 Flash Lite Preview 09 2025 ThinkingGoogle | 75.43% | Vals AIOriginal result | Not reported | Unknown |
| 56 | Claude 3.7 Sonnet 20250219 ThinkingAnthropic | 75.10% | Vals AIOriginal result | Not reported | Unknown |
| 57 | Claude Sonnet 4 20250514 ThinkingAnthropic | 74.93% | Vals AIOriginal result | Not reported | Unknown |
| 58 | Claude Opus 4.1Anthropic | 73.72% | Vals AIOriginal result | Not reported | Unknown |
| 59 | GPT-5.4 nanoOpenAI | 73.58% | high reasoningOriginal result | Not reported | Closed |
| 60 | Claude Opus 4Anthropic | 73.31% | Vals AIOriginal result | Not reported | Unknown |
| 61 | Grok 4 Fast (Reasoning)xAI | 72.78% | Vals AIOriginal result | Not reported | Closed |
| 62 | Grok 4.1 Fast (Reasoning)xAI | 72.66% | Vals AIOriginal result | Not reported | Closed |
| 63 | Gemini 2.5 Flash Lite Preview 09 2025Google | 72.54% | Vals AIOriginal result | Not reported | Unknown |
| 64 | Claude Sonnet 4Anthropic | 72.39% | Vals AIOriginal result | Not reported | Unknown |
| 65 | GPT-4.1OpenAI | 72.39% | high reasoningOriginal result | Not reported | Closed |
| 66 | Gemini 2.5 Flash Preview 04 17 ThinkingGoogle | 71.92% | Vals AIOriginal result | Not reported | Unknown |
| 67 | Llama4 Maverick Instruct BasicFireworks AI | 71.69% | Vals AIOriginal result | Not reported | Unknown |
| 68 | Claude 3.7 SonnetAnthropic | 71.52% | Vals AIOriginal result | Not reported | Unknown |
| 69 | GPT-5 nanoOpenAI | 70.94% | high reasoningOriginal result | Not reported | Closed |
| 70 | GPT-4.1 miniOpenAI | 70.54% | high reasoningOriginal result | Not reported | Closed |
| 71 | Gemini 2.0 Flash 001Google | 69.79% | Vals AIOriginal result | Not reported | Unknown |
| 72 | Claude 3.5 SonnetAnthropic | 68.80% | Vals AIOriginal result | Not reported | Closed |
| 73 | Gemini 2.5 Flash Preview 04 17Google | 68.00% | Vals AIOriginal result | Not reported | Unknown |
| 74 | Mistral Large 2512Mistral AI | 66.19% | Vals AIOriginal result | Not reported | Unknown |
| 75 | Gemini 1.5 Pro 002Google | 65.51% | Vals AIOriginal result | Not reported | Unknown |
| 76 | Magistral Small 2509Mistral AI | 65.20% | Vals AIOriginal result | Not reported | Unknown |
| 77 | Magistral Medium 2509Mistral AI | 64.57% | Vals AIOriginal result | Not reported | Unknown |
| 78 | GPT-4oOpenAI | 64.01% | high reasoningOriginal result | Not reported | Closed |
| 79 | Grok 4.1 Fast Non ReasoningxAI | 63.70% | Vals AIOriginal result | Not reported | Unknown |
| 80 | Grok 4 Fast Non ReasoningxAI | 63.41% | Vals AIOriginal result | Not reported | Unknown |
| 81 | Mistral Medium 2505Mistral AI | 62.97% | Vals AIOriginal result | Not reported | Unknown |
| 82 | GPT-4oOpenAI | 62.16% | high reasoningOriginal result | Not reported | Closed |
| 83 | Grok 4.5xAI | 61.79% | high reasoningOriginal result | Not reported | Closed |
| 84 | Mistral Small 2503Mistral AI | 60.08% | Vals AIOriginal result | Not reported | Unknown |
| 85 | Meta Llama Llama 4 Scout 17B 16E InstructTogether AI | 58.75% | Vals AIOriginal result | Not reported | Unknown |
| 86 | Grok 2 Vision 1212xAI | 57.25% | Vals AIOriginal result | Not reported | Unknown |
| 87 | Gemini 1.5 Flash 002Google | 57.19% | Vals AIOriginal result | Not reported | Unknown |
| 88 | GPT-4o miniOpenAI | 56.56% | high reasoningOriginal result | Not reported | Closed |
| 89 | GPT-4.1 nanoOpenAI | 55.05% | high reasoningOriginal result | Not reported | Closed |
| 90 | Meta Llama Llama 3.2 90B Vision Instruct TurboTogether AI | 48.06% | Vals AIOriginal result | Not reported | Unknown |
| 91 | Claude Haiku 4.5 ThinkingAnthropic | 46.07% | Vals AIOriginal result | Not reported | Closed |
| 92 | Meta Llama Llama 3.2 11B Vision Instruct TurboTogether AI | 38.82% | Vals AIOriginal result | Not reported | Unknown |
| 93 | Qwen3.5 Plus ThinkingAlibaba | 22.77% | Vals AIOriginal result | Not reported | Unknown |
Among the reported MMMU rows, Qwen3.6 Plus is first at 86.0%. The third row is 3.1 points behind. The broader top-10 range is 40.0 points, so the table still separates the published systems.
11 models have been evaluated on MMMU. The benchmark falls in the Multimodal & Grounded category. MMMU is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About MMMU
Year
2024
Tasks
Multimodal academic reasoning
Format
Image + text question answering
Difficulty
Frontier multimodal
MMMU is the base benchmark family behind later MMMU-Pro variants. It measures whether a model can answer expert-style questions that require combining visual understanding with domain knowledge and reasoning.
Freshness and provenance
Version
MMMU 2024
Refresh cadence
Annual
Staleness state
Refreshing
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does MMMU measure?
A broad multimodal reasoning benchmark spanning charts, diagrams, tables, and academic visual question answering.
Which model scores highest on MMMU?
Qwen3.6 Plus by Alibaba currently leads with a score of 86.0% on MMMU.
How many models are evaluated on MMMU?
11 AI models have published results on MMMU in the BenchLM catalog.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 5,500+ readers.
One email each week. Unsubscribe anytime.