Skip to main content
BenchLM
MIXED GLOBAL + REGIONAL

Korean Benchmarks Leaderboard

Data verified

How do global frontier models stack up against regional Korean models on domestic tasks? This leaderboard ranks all models based exclusively on Korean benchmark performance like KMMLU, KMMLU-Hard, CLIcK, and KoBALT.

Solar Open currently leads the cross-market Korean view with an average score of 87.1.

This is the right page for deciding whether Korean-market specialists are actually outperforming global frontier models on Korean-native evaluation, rather than just inside a regional-only pool.

RankModelTypeKMMLUKMMLU-HardAvg Score
#1Solar Open
Upstage
GLOBAL87.1
#2A.X
SK Telecom
GLOBAL81.7
#3Solar
Upstage
GLOBAL79.2
#4K-EXAONE 2.0
LG AI Research
GLOBAL76.7
#5Kanana 2
Kakao
GLOBAL43.32%43.3
#6Kanana 2
Kakao
GLOBAL42.79%42.8

What these rows mean

KMMLU: Measures massive multitask language understanding on 45 Korean expert-level subjects.

KMMLU-Hard: A computationally heavier slice focusing on complex Korean reasoning where models struggle most.

How to interpret the crossover

While global frontier models like GPT-5 and Claude lead in general reasoning, models like HyperClova X and Exaone are explicitly trained on high-quality Korean corpora. This leaderboard tracks the crossover points between sheer model scale and regional specialization.

View regional-only Korean LLMs

Korean benchmark updates

Get leaderboard shifts when Korean benchmark scores change for either regional or global models.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.

Recommended next step

If the mixed leaderboard shows a Korean-market model winning on your target rows, open its model page next and inspect the full score breakdown before choosing it over a global default.