# Korean Massive Multitask Language Understanding (KMMLU)

> Evaluates Korean expert-level knowledge across 45 subjects. 20% of questions require Korean cultural context.

Canonical page: https://benchlm.ai/benchmarks/kmmlu

- Category: [Korean Benchmarks](/leaderboards/korean-benchmarks)
- Last updated: September 29, 2026

## About KMMLU

- Year: 2024
- Tasks: 35,030 questions
- Format: Multiple choice questions
- Difficulty: Elementary to professional level in Korean
- Paper: [KMMLU: Measuring Massive Multitask Language Understanding in Korean](https://arxiv.org/abs/2402.11548)

Tests human-level understanding and reasoning in the Korean language across diverse subjects.

KMMLU is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (2 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Kanana-2 3B Instruct](/models/kanana-2-3b-instruct) | Kakao | 43.3% |
| 2 | [Kanana-2 1.3B Instruct](/models/kanana-2-1-3b-instruct) | Kakao | 42.8% |

## FAQ

### What does KMMLU measure?

Evaluates Korean expert-level knowledge across 45 subjects. 20% of questions require Korean cultural context.

### Which model scores highest on KMMLU?

Kanana-2 3B Instruct by Kakao currently leads with a score of 43.3% on KMMLU.

### How many models are evaluated on KMMLU?

2 AI models have been evaluated on KMMLU on BenchLM.

### Does KMMLU affect BenchLM's overall score?

Not directly. KMMLU is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on KMMLU

- [Kanana-2 3B Instruct vs Kanana-2 1.3B Instruct](/compare/kanana-2-1-3b-instruct-vs-kanana-2-3b-instruct)
