KMMLU-Hard
We show this table for reference; we do not rank on it.
A filtered hard subset of KMMLU containing ~5,000 questions that most models get wrong.
About KMMLU-Hard
Year
2025
Tasks
~5,000 questions
Format
Multiple choice questions
Difficulty
Advanced Korean reasoning
Provides strong signals for advanced frontier models attempting reasoning in Korean.
Freshness and provenance
Version
KMMLU-Hard 2025
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does KMMLU-Hard measure?
A filtered hard subset of KMMLU containing ~5,000 questions that most models get wrong.
Which model scores highest on KMMLU-Hard?
No models have been evaluated on KMMLU-Hard yet.
How many models are evaluated on KMMLU-Hard?
0 AI models have been evaluated on KMMLU-Hard on BenchLM.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.