Korean Massive Multitask Language Understanding (KMMLU)
We show this table for reference; we do not rank on it.
Evaluates Korean expert-level knowledge across 45 subjects. 20% of questions require Korean cultural context.
Benchmark score on KMMLU — September 29, 2026
We compile the KMMLU rows from provider self-reports. Kanana-2 3B Instruct leads the table at 43.3%, followed by Kanana-2 1.3B Instruct (42.8%). We do not use these results to rank models overall.
Kanana-2 3B Instruct
Kakao
Kanana-2 1.3B Instruct
Kakao
2 modelsKoreanKorean-language benchmarkRefreshingDisplay onlyUpdated September 29, 2026
Benchmark score table (2 models)
ScoreAbout KMMLU
Year
2024
Tasks
35,030 questions
Format
Multiple choice questions
Difficulty
Elementary to professional level in Korean
Tests human-level understanding and reasoning in the Korean language across diverse subjects.
Freshness and provenance
Version
KMMLU 2024
Refresh cadence
Annual
Staleness state
Refreshing
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does KMMLU measure?
Evaluates Korean expert-level knowledge across 45 subjects. 20% of questions require Korean cultural context.
Which model scores highest on KMMLU?
Kanana-2 3B Instruct by Kakao currently leads with a score of 43.3% on KMMLU.
How many models are evaluated on KMMLU?
2 AI models have been evaluated on KMMLU on BenchLM.
Compare top models on KMMLU
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.