Skip to main content

Benchmark profile

Multi-task Indic Language Understanding Benchmark (MILU)

Culturally grounded knowledge comprehension across ten Indic languages and English.

Data verified

Benchmark score on MILU — July 24, 2026

BenchLM mirrors the published score view for MILU. Claude Opus 5 leads the public snapshot at 92.1%. BenchLM does not use these results to rank models overall.

1 modelMultilingualRefreshingDisplay onlyUpdated July 24, 2026

Benchmark score table (1 model)

Score
1
Claude Opus 5Anthropic · Closed
92.1%

About MILU

Year

2024

Tasks

Knowledge tasks across 11 languages

Format

Average accuracy

Difficulty

Multilingual Indic knowledge

Anthropic reports average accuracy over five max-effort trials without tools or a custom system prompt.

BenchLM freshness & provenance

Version

MILU 2024

Refresh cadence

Annual

Staleness state

Refreshing

Question availability

Public benchmark set

RefreshingDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does MILU measure?

Culturally grounded knowledge comprehension across ten Indic languages and English.

Which model scores highest on MILU?

Claude Opus 5 by Anthropic currently leads with a score of 92.1% on MILU.

How many models are evaluated on MILU?

1 AI models have been evaluated on MILU on BenchLM.

Last updated: July 24, 2026 · BenchLM version MILU 2024