Global MMLU (GMMLU)
We show this table for reference; we do not rank on it.
MMLU-style knowledge evaluation across 42 high- and low-resource languages.
Benchmark score on GMMLU — September 27, 2026
We compile the GMMLU rows from provider self-reports. Claude Opus 5.5 leads the table at 94.3%, followed by Claude Opus 5 (92.5%). We do not use these results to rank models overall.
Claude Opus 5.5
Anthropic
Claude Opus 5
Anthropic
2 modelsMultilingualRefreshingDisplay onlyUpdated September 27, 2026
Benchmark score table (2 models)
ScoreAbout GMMLU
Year
2024
Tasks
Knowledge questions across 42 languages
Format
Average accuracy
Difficulty
Multilingual knowledge
Anthropic reports average accuracy from one max-effort trial without tools or a custom system prompt.
Freshness and provenance
Version
GMMLU 2024
Refresh cadence
Annual
Staleness state
Refreshing
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does GMMLU measure?
MMLU-style knowledge evaluation across 42 high- and low-resource languages.
Which model scores highest on GMMLU?
Claude Opus 5.5 by Anthropic currently leads with a score of 94.3% on GMMLU.
How many models are evaluated on GMMLU?
2 AI models have been evaluated on GMMLU on BenchLM.
Compare top models on GMMLU
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.