Skip to main content
BenchLM

KMMLU-Hard

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

A filtered hard subset of KMMLU containing ~5,000 questions that most models get wrong.

About KMMLU-Hard

Year

2025

Tasks

~5,000 questions

Format

Multiple choice questions

Difficulty

Advanced Korean reasoning

Provides strong signals for advanced frontier models attempting reasoning in Korean.

Freshness and provenance

Version

KMMLU-Hard 2025

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does KMMLU-Hard measure?

A filtered hard subset of KMMLU containing ~5,000 questions that most models get wrong.

Which model scores highest on KMMLU-Hard?

No models have been evaluated on KMMLU-Hard yet.

How many models are evaluated on KMMLU-Hard?

0 AI models have been evaluated on KMMLU-Hard on BenchLM.

Last updated: September 29, 2026 · BenchLM version KMMLU-Hard 2025

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.