# Multi-task Indic Language Understanding Benchmark (MILU)

> Culturally grounded knowledge comprehension across ten Indic languages and English.

Canonical page: https://benchlm.ai/benchmarks/milu

- Category: [Multilingual](/multilingual)
- Last updated: September 27, 2026

## About MILU

- Year: 2024
- Tasks: Knowledge tasks across 11 languages
- Format: Average accuracy
- Difficulty: Multilingual Indic knowledge
- Paper: [MILU: A Multi-task Indic language understanding benchmark](https://arxiv.org/abs/2411.02538)

Anthropic reports average accuracy over five max-effort trials without tools or a custom system prompt.

MILU is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (2 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Claude Opus 5.5](/models/claude-opus-5-5) | Anthropic | 93.1% |
| 2 | [Claude Opus 5](/models/claude-opus-5) | Anthropic | 92.1% |

## FAQ

### What does MILU measure?

Culturally grounded knowledge comprehension across ten Indic languages and English.

### Which model scores highest on MILU?

Claude Opus 5.5 by Anthropic currently leads with a score of 93.1% on MILU.

### How many models are evaluated on MILU?

2 AI models have been evaluated on MILU on BenchLM.

### Does MILU affect BenchLM's overall score?

Not directly. MILU is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on MILU

- [Claude Opus 5.5 vs Claude Opus 5](/compare/claude-opus-5-vs-claude-opus-5-5)
