Benchmark profile
Multi-task Indic Language Understanding Benchmark (MILU)
Culturally grounded knowledge comprehension across ten Indic languages and English.
Data verifiedBenchmark score on MILU — July 24, 2026
BenchLM mirrors the published score view for MILU. Claude Opus 5 leads the public snapshot at 92.1%. BenchLM does not use these results to rank models overall.
Benchmark score table (1 model)
ScoreAbout MILU
Year
2024
Tasks
Knowledge tasks across 11 languages
Format
Average accuracy
Difficulty
Multilingual Indic knowledge
Anthropic reports average accuracy over five max-effort trials without tools or a custom system prompt.
BenchLM freshness & provenance
Version
MILU 2024
Refresh cadence
Annual
Staleness state
Refreshing
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does MILU measure?
Culturally grounded knowledge comprehension across ten Indic languages and English.
Which model scores highest on MILU?
Claude Opus 5 by Anthropic currently leads with a score of 92.1% on MILU.
How many models are evaluated on MILU?
1 AI models have been evaluated on MILU on BenchLM.