# KMMLU-Hard

> A filtered hard subset of KMMLU containing ~5,000 questions that most models get wrong.

Canonical page: https://benchlm.ai/benchmarks/kmmluhard

- Category: [Korean Benchmarks](/leaderboards/korean-benchmarks)
- Last updated: September 29, 2026

## About KMMLU-Hard

- Year: 2025
- Tasks: ~5,000 questions
- Format: Multiple choice questions
- Difficulty: Advanced Korean reasoning
- Paper: [Evaluating LLMs on Hard Korean Queries](https://github.com/daekeun-ml/evaluate-llm-on-korean-dataset)

Provides strong signals for advanced frontier models attempting reasoning in Korean.

KMMLU-Hard is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (0 models)

Benchmark data for this page is coming soon.

## FAQ

### What does KMMLU-Hard measure?

A filtered hard subset of KMMLU containing ~5,000 questions that most models get wrong.

### Which model scores highest on KMMLU-Hard?

No models have been evaluated on KMMLU-Hard yet.

### How many models are evaluated on KMMLU-Hard?

0 AI models have been evaluated on KMMLU-Hard on BenchLM.

### Does KMMLU-Hard affect BenchLM's overall score?

Not directly. KMMLU-Hard is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
