Massive Multitask Language Understanding (MMLU)
A comprehensive multiple-choice question answering test covering 57 tasks including elementary mathematics, US history, computer science, law, and more. Tests knowledge across diverse academic subjects from high school to professional level.
Data verified 27 confirmed releases in the last 30 daysSee the free Radar BriefBenchmark score on MMLU — September 2, 2026
We mirror the published score view for MMLU. o1 leads the public snapshot at 91.8%, followed by GPT-4.1 (90.2%) and DeepSeek V4 Pro Base (90.1%). We do not use these results to rank models overall.
o1
OpenAI
GPT-4.1
OpenAI
DeepSeek V4 Pro Base
DeepSeek
8 modelsKnowledgeStaleSaturatedDisplay onlyUpdated September 2, 2026
Benchmark score table (8 models)
ScoreThe published MMLU snapshot places o1 first at 91.8%. The third row is 1.7 points behind. The broader top-10 range is 11.7 points, so the table still separates the published systems.
8 models have been evaluated on MMLU. The benchmark falls in the Knowledge category. This category carries a 12% weight in BenchLM.ai's overall scoring system. MMLU is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About MMLU
Year
2020
Tasks
57 subjects
Format
Multiple choice questions
Difficulty
Elementary to professional level
MMLU evaluates models on 57 subjects spanning humanities, social sciences, STEM, and other areas. Questions range from elementary to advanced professional level, making it a comprehensive test of world knowledge and reasoning ability.
BenchLM freshness & provenance
Version
MMLU
Refresh cadence
Static
Staleness state
Stale
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does MMLU measure?
A comprehensive multiple-choice question answering test covering 57 tasks including elementary mathematics, US history, computer science, law, and more. Tests knowledge across diverse academic subjects from high school to professional level.
Which model scores highest on MMLU?
o1 by OpenAI currently leads with a score of 91.8% on MMLU.
How many models are evaluated on MMLU?
8 AI models have been evaluated on MMLU on BenchLM.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.