Skip to main content

Benchmark profile

Massive Multitask Language Understanding (MMLU)

A comprehensive multiple-choice question answering test covering 57 tasks including elementary mathematics, US history, computer science, law, and more. Tests knowledge across diverse academic subjects from high school to professional level.

Data verified

Benchmark score on MMLU — July 20, 2026

BenchLM mirrors the published score view for MMLU. o1 leads the public snapshot at 91.8% , followed by GPT-4.1 (90.2%) and DeepSeek V4 Pro Base (90.1%). BenchLM does not use these results to rank models overall.

8 modelsKnowledgeStaleSaturatedDisplay onlyUpdated July 20, 2026

Benchmark score table (8 models)

Score
1
o1OpenAI · Closed
91.8%
2
GPT-4.1OpenAI · Closed
90.2%
3
DeepSeek V4 Pro BaseDeepSeek · Open weight
90.1%
4
DeepSeek V4 Flash BaseDeepSeek · Open weight
88.7%
5
GPT-4.1 miniOpenAI · Closed
87.5%
6
Trinity-Large-PreviewArcee AI · Open weight
87.2%
7
o3-miniOpenAI · Closed
86.9%
8
GPT-4.1 nanoOpenAI · Closed
80.1%

The published MMLU snapshot places o1 first at 91.8%. The third row is 1.7 points behind. The broader top-10 range is 11.7 points, so the table still separates the published systems.

8 models have been evaluated on MMLU. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. MMLU is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About MMLU

Year

2020

Tasks

57 subjects

Format

Multiple choice questions

Difficulty

Elementary to professional level

MMLU evaluates models on 57 subjects spanning humanities, social sciences, STEM, and other areas. Questions range from elementary to advanced professional level, making it a comprehensive test of world knowledge and reasoning ability.

BenchLM freshness & provenance

Version

MMLU

Refresh cadence

Static

Staleness state

Stale

Question availability

Public benchmark set

StaleSaturatedDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does MMLU measure?

A comprehensive multiple-choice question answering test covering 57 tasks including elementary mathematics, US history, computer science, law, and more. Tests knowledge across diverse academic subjects from high school to professional level.

Which model scores highest on MMLU?

o1 by OpenAI currently leads with a score of 91.8% on MMLU.

How many models are evaluated on MMLU?

8 AI models have been evaluated on MMLU on BenchLM.

Last updated: July 20, 2026 · BenchLM version MMLU

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.