# MMMLU

> A multilingual MMLU-style benchmark reported in provider evaluation tables.

Canonical page: https://benchlm.ai/benchmarks/mmmlu

- Category: [Knowledge](/knowledge)
- Last updated: September 27, 2026

## About MMMLU

- Year: 2026
- Tasks: Multilingual academic QA
- Format: Exact match
- Difficulty: Broad multilingual knowledge
- Paper: [MMMLU](https://huggingface.co/datasets/openai/MMMLU)

BenchLM stores MMMLU as a display-only provider-table row when exact public values are published.

MMMLU is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (5 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Interfaze Beta](/models/interfaze-beta) | Interfaze | 90.9% |
| 2 | [Qwen3.7 Max](/models/qwen3-7-max) | Alibaba | 90.3% |
| 3 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 89.0% |
| 4 | [K-EXAONE 2.0](/models/k-exaone-2-0) | LG AI Research | 86.6% |
| 5 | [Gemma 4 12B](/models/gemma-4-12b) | Google | 83.4% |

## FAQ

### What does MMMLU measure?

A multilingual MMLU-style benchmark reported in provider evaluation tables.

### Which model scores highest on MMMLU?

Interfaze Beta by Interfaze currently leads with a score of 90.9% on MMMLU.

### How many models are evaluated on MMMLU?

5 AI models have been evaluated on MMMLU on BenchLM.

### Does MMMLU affect BenchLM's overall score?

Not directly. MMMLU is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on MMMLU

- [Interfaze Beta vs Qwen3.7 Max](/compare/interfaze-beta-vs-qwen3-7-max)
- [Qwen3.7 Max vs Qwen3.7 Plus](/compare/qwen3-7-max-vs-qwen3-7-plus)
- [Qwen3.7 Plus vs K-EXAONE 2.0](/compare/k-exaone-2-0-vs-qwen3-7-plus)
- [K-EXAONE 2.0 vs Gemma 4 12B](/compare/gemma-4-12b-vs-k-exaone-2-0)
