Skip to main content
BenchLM

Multi-task Indic Language Understanding Benchmark (MILU)

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

Culturally grounded knowledge comprehension across ten Indic languages and English.

Benchmark score on MILU — September 27, 2026

We compile the MILU rows from provider self-reports. Claude Opus 5.5 leads the table at 93.1%, followed by Claude Opus 5 (92.1%). We do not use these results to rank models overall.

2 modelsMultilingualRefreshingDisplay onlyUpdated September 27, 2026

Benchmark score table (2 models)

Score
1
Claude Opus 5.5Anthropic · Closed
93.1%
2
Claude Opus 5Anthropic · Closed
92.1%

About MILU

Year

2024

Tasks

Knowledge tasks across 11 languages

Format

Average accuracy

Difficulty

Multilingual Indic knowledge

Anthropic reports average accuracy over five max-effort trials without tools or a custom system prompt.

Freshness and provenance

Version

MILU 2024

Refresh cadence

Annual

Staleness state

Refreshing

Question availability

Public benchmark set

RefreshingDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does MILU measure?

Culturally grounded knowledge comprehension across ten Indic languages and English.

Which model scores highest on MILU?

Claude Opus 5.5 by Anthropic currently leads with a score of 93.1% on MILU.

How many models are evaluated on MILU?

2 AI models have been evaluated on MILU on BenchLM.

Compare top models on MILU

Last updated: September 27, 2026 · BenchLM version MILU 2024

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.