Skip to main content
BenchLM

MMLU-Redux

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

A harder refresh of MMLU intended to keep broad knowledge evaluation useful after the original benchmark became too easy for frontier models.

Benchmark score on MMLU-Redux — September 27, 2026

We compile the MMLU-Redux rows from provider self-reports and secondary reports. Claude Opus 4.5 leads the table at 96.6%, followed by Qwen3.7 Max (95%) and Qwen3.5 397B (94.9%). We do not use these results to rank models overall.

13 modelsKnowledgeCurrentDisplay onlyUpdated September 27, 2026

Benchmark score table (13 models)

Score
1
Claude Opus 4.5Anthropic · Closed
96.6%
2
Qwen3.7 MaxAlibaba · Closed
95%
3
Qwen3.5 397BAlibaba · Open weight
94.9%
4
Qwen3.7 PlusAlibaba · Closed
94.5%
5
Qwen3.6 PlusAlibaba · Closed
94.5%
6
Qwen3.6-27BAlibaba · Open weight
93.5%
7
DeepSeek V4 Pro BaseDeepSeek · Open weight
90.8%
8
DeepSeek V4 Flash BaseDeepSeek · Open weight
89.4%
9
Ternary Bonsai 2 27BPrism ML · Open weight
89.1%
10
Mellum2-12B-A2.5B-ThinkingJetBrains · Open weight
86.2%
11
MiniCPM5-2BOpenBMB · Open weight
84.7%
12
Mellum2-12B-A2.5B-InstructJetBrains · Open weight
78.1%
13
MiniCPM5-1BOpenBMB · Open weight
70.1%

Among the reported MMLU-Redux rows, Claude Opus 4.5 is first at 96.6%. The third row is 1.7 points behind. The broader top-10 range is 10.4 points, so the table still separates the published systems.

13 models have been evaluated on MMLU-Redux. The benchmark falls in the Knowledge category. MMLU-Redux is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About MMLU-Redux

Year

2026

Tasks

Broad academic QA

Format

Multiple choice questions

Difficulty

Advanced general knowledge

MMLU-Redux is useful when MMLU itself has largely saturated. It acts as a broader knowledge sanity check with fresher or harder questions intended to preserve separation among strong general-purpose models.

Freshness and provenance

Version

MMLU-Redux 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does MMLU-Redux measure?

A harder refresh of MMLU intended to keep broad knowledge evaluation useful after the original benchmark became too easy for frontier models.

Which model scores highest on MMLU-Redux?

Claude Opus 4.5 by Anthropic currently leads with a score of 96.6% on MMLU-Redux.

How many models are evaluated on MMLU-Redux?

13 AI models have been evaluated on MMLU-Redux on BenchLM.

Last updated: September 27, 2026 · BenchLM version MMLU-Redux 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.