# Global MMLU (GMMLU)

> MMLU-style knowledge evaluation across 42 high- and low-resource languages.

Canonical page: https://benchlm.ai/benchmarks/gmmlu

- Category: [Multilingual](/multilingual)
- Last updated: September 27, 2026

## About GMMLU

- Year: 2024
- Tasks: Knowledge questions across 42 languages
- Format: Average accuracy
- Difficulty: Multilingual knowledge
- Paper: [Global MMLU: Understanding and addressing cultural and linguistic biases in multilingual evaluation](https://arxiv.org/abs/2412.03304)

Anthropic reports average accuracy from one max-effort trial without tools or a custom system prompt.

GMMLU is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (2 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 5.5](/models/claude-opus-5-5) | Anthropic | 94.3% |
| 2 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 92.5% |

## FAQ

### What does GMMLU measure?

MMLU-style knowledge evaluation across 42 high- and low-resource languages.

### Which model scores highest on GMMLU?

Claude Opus 5.5 by Anthropic currently leads with a score of 94.3% on GMMLU.

### How many models are evaluated on GMMLU?

2 AI models have been evaluated on GMMLU on BenchLM.

### Does GMMLU affect BenchLM's overall score?

Not directly. GMMLU is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on GMMLU

- [Claude Opus 5.5 vs Claude Opus 5](/compare/claude-opus-5-vs-claude-opus-5-5)
