Benchmark profile
Massive Multitask Language Understanding Professional (MMLU-Pro)
An enhanced version of MMLU with 10 answer choices instead of 4, featuring more reasoning-focused questions that better differentiate frontier models.
Data verifiedTop models on MMLU-Pro — July 20, 2026
As of July 20, 2026, Qwen3.7 Max leads the MMLU-Pro leaderboard with 89.6% , followed by Claude Opus 4.5 (89.5%) and Qwen3.7 Plus (88.5%).
Qwen3.7 Max
Alibaba
qwen3-7-max
Claude Opus 4.5
Anthropic
claude-opus-4-5
Qwen3.7 Plus
Alibaba
qwen3-7-plus
Leaderboard (42 models)
ScoreAccording to BenchLM.ai, Qwen3.7 Max leads the MMLU-Pro benchmark with a score of 89.6%, followed by Claude Opus 4.5 (89.5%) and Qwen3.7 Plus (88.5%). The top models are clustered within 1.1 points, suggesting this benchmark is nearing saturation for frontier models.
42 models have been evaluated on MMLU-Pro. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. Within that category, MMLU-Pro contributes 30% of the category score, so strong performance here directly affects a model's overall ranking.
About MMLU-Pro
Year
2024
Tasks
Multiple subjects
Format
10-way multiple choice
Difficulty
Professional level
MMLU-Pro increases the number of choices from 4 to 10 and integrates more reasoning-focused problems, reducing the chance of correct guessing and better evaluating true understanding. It serves as a more robust discriminator of model capabilities.
BenchLM freshness & provenance
Version
MMLU-Pro
Refresh cadence
Static
Staleness state
Refreshing
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does MMLU-Pro measure?
An enhanced version of MMLU with 10 answer choices instead of 4, featuring more reasoning-focused questions that better differentiate frontier models.
Which model scores highest on MMLU-Pro?
Qwen3.7 Max by Alibaba currently leads with a score of 89.6% on MMLU-Pro.
How many models are evaluated on MMLU-Pro?
42 AI models have been evaluated on MMLU-Pro on BenchLM.
Compare Top Models on MMLU-Pro
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.