# Massive Multi-discipline Multimodal Understanding Pro (MMMU-Pro)

> A harder multimodal benchmark for frontier models that combines text with images, diagrams, charts, and academic visual reasoning tasks.

Canonical page: https://benchlm.ai/benchmarks/mmmu-pro

- Category: [Multimodal & Grounded](/multimodal-grounded)
- Last updated: September 15, 2026

## About MMMU-Pro

- Year: 2024
- Tasks: Multimodal academic reasoning
- Format: Image + text question answering
- Difficulty: Frontier multimodal
- Paper: [MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark](https://arxiv.org/abs/2409.02813)

MMMU-Pro extends the original MMMU setup with more difficult multimodal questions and stronger separation at the top end of the model market.

MMMU-Pro is currently weighted in BenchLM's scoring formula. The Multimodal & Grounded category carries 12% of the overall score, and MMMU-Pro contributes 40% of that category score.

## Leaderboard (40 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Gemini 3.1 Pro](/models/gemini-3-1-pro) | Google | 83.9% |
| 2 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 83.6% |
| 3 | [GPT-5.6 Sol](/models/gpt-5-6-sol) | OpenAI | 83% |
| 4 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 82.3% |
| 5 | [Kimi K3](/models/kimi-k3) | Moonshot AI | 81.6% |
| 6 | [Seed 2.1 Pro](/models/seed-2-1-pro) | ByteDance | 81.6% |
| 7 | [GPT-5.5](/models/gpt-5-5) | OpenAI | 81.2% |
| 8 | [GPT-5.4](/models/gpt-5-4) | OpenAI | 81.2% |
| 9 | [Gemini 3 Pro](/models/gemini-3-pro) | Google | 81% |
| 10 | [GPT-5.6 Terra](/models/gpt-5-6-terra) | OpenAI | 80.7% |
| 11 | [Muse Spark](/models/muse-spark) | Meta | 80.4% |
| 12 | [Seed 2.1 Turbo](/models/seed-2-1-turbo) | ByteDance | 80.1% |
| 13 | [GPT-5.2](/models/gpt-5-2) | OpenAI | 79.5% |
| 14 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 79.4% |
| 15 | [dots3-note Preview](/models/dots3-note-preview) | Dots Studio | 79.1% |
| 16 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 79% |
| 17 | [Qwen3.5 397B](/models/qwen3-5-397b) | Alibaba | 79% |
| 18 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 78.8% |
| 19 | [Kimi K2.5 (Reasoning)](/models/kimi-k2-5-reasoning) | Moonshot AI | 78.5% |
| 20 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 78.5% |
| 21 | [GPT-5.6 Luna](/models/gpt-5-6-luna) | OpenAI | 78.4% |
| 22 | [MiniMax M3](/models/minimax-m3) | MiniMax | 78.1% |
| 23 | [Grok 4.3](/models/grok-4-3) | xAI | 78.1% |
| 24 | [MiMo-V2.5](/models/mimo-v2-5) | Xiaomi | 77.9% |
| 25 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 77.3% |
| 26 | [Gemma 4 31B](/models/gemma-4-31b) | Google | 76.9% |
| 27 | [GPT-5.4 mini](/models/gpt-5-4-mini) | OpenAI | 76.6% |
| 28 | [Qwen3.6-27B](/models/qwen3-6-27b) | Alibaba | 75.8% |
| 29 | [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) | Alibaba | 75.3% |
| 30 | [Grok 4.20](/models/grok-4-20-beta) | xAI | 75.2% |
| 31 | [Inkling-Small](/models/inkling-small) | Thinking Machines Lab | 74% |
| 32 | [Muse Glimmer 30B](/models/muse-glimmer-30b) | Meta | 74% |
| 33 | [Gemma 4 26B A4B](/models/gemma-4-26b-a4b) | Google | 73.8% |
| 34 | [Inkling](/models/inkling) | Thinking Machines Lab | 73.5% |
| 35 | [Interfaze Beta](/models/interfaze-beta) | Interfaze | 71.1% |
| 36 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 70.6% |
| 37 | [Gemma 4 12B](/models/gemma-4-12b) | Google | 69.1% |
| 38 | [GPT-5.4 nano](/models/gpt-5-4-nano) | OpenAI | 66.1% |
| 39 | [Command A+](/models/command-a-plus) | Cohere | 63% |
| 40 | [LFM2.5-VL-3B](/models/lfm2-5-vl-3b) | LiquidAI | 30.5% |

## FAQ

### What does MMMU-Pro measure?

A harder multimodal benchmark for frontier models that combines text with images, diagrams, charts, and academic visual reasoning tasks.

### Which model scores highest on MMMU-Pro?

Gemini 3.1 Pro by Google currently leads with a score of 83.9% on MMMU-Pro.

### How many models are evaluated on MMMU-Pro?

40 AI models have been evaluated on MMMU-Pro on BenchLM.

### Does MMMU-Pro affect BenchLM's overall score?

Yes. MMMU-Pro is a weighted benchmark inside the Multimodal & Grounded category, which carries 12% of BenchLM's overall score. MMMU-Pro itself contributes 40% of that category score.

## Compare Top Models on MMMU-Pro

- [Gemini 3.1 Pro vs Gemini 3.5 Flash](/compare/gemini-3-1-pro-vs-gemini-3-5-flash)
- [Gemini 3.5 Flash vs GPT-5.6 Sol](/compare/gemini-3-5-flash-vs-gpt-5-6-sol)
- [GPT-5.6 Sol vs Qwen3.8 Max](/compare/gpt-5-6-sol-vs-qwen3-8-max)
- [Qwen3.8 Max vs Kimi K3](/compare/kimi-k3-vs-qwen3-8-max)
