# Massive Multitask Language Understanding Professional (MMLU-Pro)

> An enhanced version of MMLU with 10 answer choices instead of 4, featuring more reasoning-focused questions that better differentiate frontier models.

Canonical page: https://benchlm.ai/benchmarks/mmlu-pro

- Category: [Knowledge](/knowledge)
- Last updated: September 10, 2026

## About MMLU-Pro

- Year: 2024
- Tasks: Multiple subjects
- Format: 10-way multiple choice
- Difficulty: Professional level
- Paper: [MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark](https://arxiv.org/abs/2406.01574)

MMLU-Pro increases the number of choices from 4 to 10 and integrates more reasoning-focused problems, reducing the chance of correct guessing and better evaluating true understanding. It serves as a more robust discriminator of model capabilities.

MMLU-Pro is currently weighted in BenchLM's scoring formula. The Knowledge category carries 12% of the overall score, and MMLU-Pro contributes 20% of that category score.

## Leaderboard (46 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Qwen3.7 Max](/models/qwen3-7-max) | Alibaba | 89.6% |
| 2 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 89.5% |
| 3 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 88.5% |
| 4 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 88.5% |
| 5 | [Qwen3.5 397B](/models/qwen3-5-397b) | Alibaba | 87.8% |
| 6 | [DeepSeek V4 Pro 0813](/models/deepseek-v4-pro-0813) | DeepSeek | 87.5% |
| 7 | [Kimi K2.5 (Reasoning)](/models/kimi-k2-5-reasoning) | Moonshot AI | 87.1% |
| 8 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 87.1% |
| 9 | [Nemotron 3 Ultra](/models/nemotron-3-ultra) | NVIDIA | 86.8% |
| 10 | [Qwen3.5-122B-A10B](/models/qwen3-5-122b-a10b) | Alibaba | 86.7% |
| 11 | [Solar Pro 4](/models/solar-pro-4) | Upstage | 86.3% |
| 12 | [Qwen3.6-27B](/models/qwen3-6-27b) | Alibaba | 86.2% |
| 13 | [DeepSeek V4 Flash 0731](/models/deepseek-v4-flash-0731) | DeepSeek | 86.2% |
| 14 | [Solar Open 2](/models/solar-open-2) | Upstage | 86.2% |
| 15 | [Qwen3.5-27B](/models/qwen3-5-27b) | Alibaba | 86.1% |
| 16 | [GLM-5](/models/glm-5) | Z.AI | 85.7% |
| 17 | [Qwen3.5-35B-A3B](/models/qwen3-5-35b-a3b) | Alibaba | 85.3% |
| 18 | [Gemma 4 31B](/models/gemma-4-31b) | Google | 85.2% |
| 19 | [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) | Alibaba | 85.2% |
| 20 | [MAI-Thinking-1](/models/mai-thinking-1) | Microsoft | 85% |
| 21 | [MiMo-V2-Flash](/models/mimo-v2-flash) | Xiaomi | 84.9% |
| 22 | [GLM-4.7](/models/glm-4-7) | Z.AI | 84.3% |
| 23 | [K-EXAONE 2.0](/models/k-exaone-2-0) | LG AI Research | 83.5% |
| 24 | [Qwen3 235B 2507](/models/qwen3-235b-2507) | Alibaba | 83% |
| 25 | [Gemma 4 26B A4B](/models/gemma-4-26b-a4b) | Google | 82.6% |
| 26 | [Claude Opus 4.6](/models/claude-opus-4-6) | Anthropic | 82% |
| 27 | [Exaone 4.0 32B](/models/exaone-4-0-32b) | LG AI Research | 81.8% |
| 28 | [Nemotron 3.5 Lightning 30B A3B NVFP4](/models/nemotron-3-5-lightning-30b-a3b-nvfp4) | NVIDIA | 81.6% |
| 29 | [LongCat-Flash-Lite-Sparse](/models/longcat-flash-lite-sparse) | Meituan | 79.2% |
| 30 | [Claude Sonnet 4.6](/models/claude-sonnet-4-6) | Anthropic | 79.2% |
| 31 | [Granite 4.2 30B](/models/granite-4-2-30b) | IBM | 77.6% |
| 32 | [Nemotron 3 Nano Omni 30B A3B](/models/nemotron-3-nano-omni-30b-a3b) | NVIDIA | 77.3% |
| 33 | [Gemma 4 12B](/models/gemma-4-12b) | Google | 77.2% |
| 34 | [DeepSeek V3](/models/deepseek-v3) | DeepSeek | 75.9% |
| 35 | [Celeris-1](/models/celeris-1) | Celeris | 75.9% |
| 36 | [ZAYA1-8B](/models/zaya1-8b) | Zyphra | 74.2% |
| 37 | [Granite 4.2 8B](/models/granite-4-2-8b) | IBM | 74.0% |
| 38 | [MiniCPM5-2B](/models/minicpm5-2b) | OpenBMB | 70.8% |
| 39 | [Gemma 4 E4B](/models/gemma-4-e4b) | Google | 69.4% |
| 40 | [ZAYA1-74B-Preview](/models/zaya1-74b-preview) | Zyphra | 68.1% |
| 41 | [Granite 4.2 3B](/models/granite-4-2-3b) | IBM | 67.8% |
| 42 | [Gemma 4 E2B](/models/gemma-4-e2b) | Google | 60% |
| 43 | [Soofi S 30B-A3B](/models/soofi-s-30b-a3b) | Soofi Project | 51.4% |
| 44 | [MiniCPM5-1B](/models/minicpm5-1b) | OpenBMB | 48.9% |
| 45 | [LFM2.5-230M](/models/lfm2-5-230m) | LiquidAI | 20.3% |
| 46 | [LFM2.5-VL-450M](/models/lfm2-5-vl-450m) | LiquidAI | 19.3% |

## FAQ

### What does MMLU-Pro measure?

An enhanced version of MMLU with 10 answer choices instead of 4, featuring more reasoning-focused questions that better differentiate frontier models.

### Which model scores highest on MMLU-Pro?

Qwen3.7 Max by Alibaba currently leads with a score of 89.6% on MMLU-Pro.

### How many models are evaluated on MMLU-Pro?

46 AI models have been evaluated on MMLU-Pro on BenchLM.

### Does MMLU-Pro affect BenchLM's overall score?

Yes. MMLU-Pro is a weighted benchmark inside the Knowledge category, which carries 12% of BenchLM's overall score. MMLU-Pro itself contributes 20% of that category score.

## Compare Top Models on MMLU-Pro

- [Qwen3.7 Max vs Claude Opus 4.5](/compare/claude-opus-4-5-vs-qwen3-7-max)
- [Claude Opus 4.5 vs Qwen3.7 Plus](/compare/claude-opus-4-5-vs-qwen3-7-plus)
- [Qwen3.7 Plus vs Qwen3.6 Plus](/compare/qwen3-6-plus-vs-qwen3-7-plus)
- [Qwen3.6 Plus vs Qwen3.5 397B](/compare/qwen3-5-397b-vs-qwen3-6-plus)

## Related Reading

- [MMLU-Pro benchmark explainer](/blog/posts/mmlu-vs-mmlu-pro)
