Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes

Massive Multitask Language Understanding Professional (MMLU-Pro)

An enhanced version of MMLU with 10 answer choices instead of 4, featuring more reasoning-focused questions that better differentiate frontier models.

Data verified 31 confirmed releases in the last 30 daysSee provider release alerts

Top models on MMLU-Pro — September 10, 2026

As of September 10, 2026, Qwen3.7 Max leads the MMLU-Pro leaderboard with 89.6% , followed by Claude Opus 4.5 (89.5%) and Qwen3.7 Plus (88.5%).

52 modelsKnowledge20% of category scoreRefreshingUpdated September 10, 2026

Leaderboard (52 models)

Score
1
Qwen3.7 MaxAlibaba · Closed
89.6%
2
Claude Opus 4.5Anthropic · Closed
89.5%
3
Qwen3.7 PlusAlibaba · Closed
88.5%
4
Qwen3.6 PlusAlibaba · Closed
88.5%
5
Qwen3.5 397BAlibaba · Open weight
87.8%
6
DeepSeek V4 Pro 0813DeepSeek · Closed
87.5%
7
Kimi K2.5Moonshot AI · Open weight
87.1%
8
Kimi K2.5 (Reasoning)Moonshot AI · Closed
87.1%
9
DeepSeek V4 Pro (High)DeepSeek · Open weight
87.1%
10
Nemotron 3 UltraNVIDIA · Open weight
86.8%
11
Qwen3.5-122B-A10BAlibaba · Open weight
86.7%
12
DeepSeek V4 Flash (High)DeepSeek · Closed
86.4%
13
Solar Pro 4Upstage · Closed
86.3%
14
Qwen3.6-27BAlibaba · Open weight
86.2%
15
DeepSeek V4 Flash 0731DeepSeek · Closed
86.2%
16
Solar Open 2Upstage · Open weight
86.2%
17
Qwen3.5-27BAlibaba · Open weight
86.1%
18
GLM-5Z.AI · Open weight
85.7%
19
Qwen3.5-35B-A3BAlibaba · Open weight
85.3%
20
Gemma 4 31BGoogle · Open weight
85.2%
21
Qwen3.6-35B-A3BAlibaba · Open weight
85.2%
22
MAI-Thinking-1Microsoft · Closed
85%
23
MiMo-V2-FlashXiaomi · Open weight
84.9%
24
GLM-4.7Z.AI · Open weight
84.3%
25
K-EXAONE 2.0LG AI Research · Open weight
83.5%
26
Qwen3 235B 2507Alibaba · Open weight
83%
27
DeepSeek V4 FlashDeepSeek · Closed
83%
28
DeepSeek V4 ProDeepSeek · Open weight
82.9%
29
Gemma 4 26B A4BGoogle · Open weight
82.6%
30
Claude Opus 4.6Anthropic · Closed
82%
31
Exaone 4.0 32BLG AI Research · Open weight
81.8%
32
81.6%
33
LongCat-Flash-Lite-SparseMeituan · Open weight
79.2%
34
Claude Sonnet 4.6Anthropic · Closed
79.2%
35
Granite 4.2 30BIBM · Open weight
77.6%
36
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
77.3%
37
Gemma 4 12BGoogle · Open weight
77.2%
38
Celeris-1Celeris · Closed
75.9%
39
DeepSeek V3DeepSeek · Open weight
75.9%
40
ZAYA1-8BZyphra · Open weight
74.2%
41
Granite 4.2 8BIBM · Open weight
74.0%
42
DeepSeek V4 Pro BaseDeepSeek · Open weight
73.5%
43
MiniCPM5-2BOpenBMB · Open weight
70.8%
44
Gemma 4 E4BGoogle · Open weight
69.4%
45
DeepSeek V4 Flash BaseDeepSeek · Open weight
68.3%
46
ZAYA1-74B-PreviewZyphra · Open weight
68.1%
47
Granite 4.2 3BIBM · Open weight
67.8%
48
Gemma 4 E2BGoogle · Open weight
60%
49
Soofi S 30B-A3BSoofi Project · Open weight
51.4%
50
MiniCPM5-1BOpenBMB · Open weight
48.9%
51
LFM2.5-230MLiquidAI · Open weight
20.3%
52
LFM2.5-VL-450MLiquidAI · Open weight
19.3%

According to BenchLM.ai, Qwen3.7 Max leads the MMLU-Pro benchmark with a score of 89.6%, followed by Claude Opus 4.5 (89.5%) and Qwen3.7 Plus (88.5%). The top models are clustered within 1.1 points, suggesting this benchmark is nearing saturation for frontier models.

52 models have been evaluated on MMLU-Pro. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. Within that category, MMLU-Pro contributes 20% of the category score, so strong performance here directly affects a model's overall ranking.

About MMLU-Pro

Year

2024

Tasks

Multiple subjects

Format

10-way multiple choice

Difficulty

Professional level

MMLU-Pro increases the number of choices from 4 to 10 and integrates more reasoning-focused problems, reducing the chance of correct guessing and better evaluating true understanding. It serves as a more robust discriminator of model capabilities.

BenchLM freshness & provenance

Version

MMLU-Pro

Refresh cadence

Static

Staleness state

Refreshing

Question availability

Public benchmark set

Refreshing

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does MMLU-Pro measure?

An enhanced version of MMLU with 10 answer choices instead of 4, featuring more reasoning-focused questions that better differentiate frontier models.

Which model scores highest on MMLU-Pro?

Qwen3.7 Max by Alibaba currently leads with a score of 89.6% on MMLU-Pro.

How many models are evaluated on MMLU-Pro?

52 AI models have been evaluated on MMLU-Pro on BenchLM.

Last updated: September 10, 2026 · BenchLM version MMLU-Pro

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.