Best LLMs for Multilingual — July 2026 Leaderboard
As of July 2026, the top multilingual model on the BenchLM leaderboard is Qwen3.7 Max with a weighted multilingual score of 100.
Data refreshed:
Performance across multiple languages
Decision lens: use provisional-ranked mode for broader public evidence and verified-ranked mode for source-only comparisons. A model can move between views as evidence coverage changes.
- Data refreshed
- July 21, 2026
- Provisional-ranked
- 13 of 290 models
- Verified-ranked
- 12 of 290 models
- Weighted evidence
- 1 of 2 benchmarks
2 tracked benchmarks
MGSM, MMLU-ProX
Evidence set: MGSM, MMLU-ProX
Best Multilingual picks
BenchLM summaries for multilingual plus the practical tradeoffs users check next: open weights, price, speed, latency, and context.
Multilingual Leaderboard
Primary score: weighted multilingual score. Higher values rank first. Use the Show metric control to change the value shown in each row.
Switch between provisional-ranked and verified-ranked modes to compare the broader public dataset with sourced-only rankings.
Filters
1 Qwen3.7 Max Alibaba | 100% | 78 | — | 87% |
2 Claude Opus 4.5 Anthropic | 82.9% | 60 | 90%P | 85.7% |
3 Qwen3.7 Plus Alibaba | 78.9% | 74 | — | 85.4% |
4 Qwen3.6 Plus Alibaba | 69.7% | 64 | — | 84.7% |
5 | 69.7% | 59 | 82%P | 84.7% |
6 | 48.7% | 64 | 84%P | 83.1% |
7 | 47.4% | 55 | — | 83% |
| 38.2% | 56 | 83%P | 82.3% | |
9 | 36.8% | 55 | — | 82.2% |
10 | 36.8% | 54 | — | 82.2% |
11 | 21.1% | 49 | — | 81% |
12 | 1% | Est.53 | 63%P | 79.4% |
13 Claude Mythos 5 Anthropic | 0% | 87 | — | 92.7%P |
Top AI Models for Multilingual — July 2026
As of July 2026, Qwen3.7 Max leads the provisional multilingual leaderboard with a score of 100.0%, followed by Claude Opus 4.5 (82.9%) and Qwen3.7 Plus (78.9%). BenchLM is currently showing 13 provisional-ranked models and 12 verified-ranked models in this category.
What changed
Claude Mythos Preview leads multilingual with the most consistent cross-language scores.
GPT-5.4 close second, strong on MMLU-ProX across all tested languages.
Claude Opus 4.6 holds #3, with particularly strong MGSM performance.
How to choose
Top models by benchmark
Broad multilingual professional benchmark across many languages(100% of category score)
Score in Context
What these scores mean
Multilingual carries a 7% weight in overall scoring. The weighted score blends MGSM (multilingual math reasoning) and MMLU-ProX (cross-language professional knowledge). This category reveals how well model capabilities transfer beyond English, where most training data is concentrated.
Known limitations
Only two benchmarks cover this category, which limits the signal. MGSM tests math reasoning specifically, not general language quality. Languages tested are limited — low-resource languages remain untested. A model scoring well here may still struggle with less common languages or dialects.
How we weight
Multilingual carries a 7% weight in BenchLM.ai's overall scoring. Cross-language performance reveals how well model capabilities transfer beyond English. See the multilingual leaderboard or compare with knowledge benchmarks.
Leaderboards exclude benchmark rows that BenchLM generated from other scores or cloned from reference models. When a weighted benchmark is missing after that filter, the category falls back to the remaining trustworthy public rows instead of filling the gap with synthetic values.
The full scoring rules, freshness handling, and runtime/pricing caveats live on the BenchLM methodology page.
Scroll horizontally to read the full evidence ledger.
| Benchmark | Weight | Status | Description |
|---|---|---|---|
| MGSM | — | Display only | Grade school math problems translated into 10 diverse languages plus English |
| MMLU-ProX | 100% | Weighted | Broad multilingual professional benchmark across many languages |
About Multilingual Benchmarks
Grade school math problems translated into 10 diverse languages plus English
Common questions
What is the best LLM for multilingual tasks?
The top multilingual LLMs are ranked by benchmarks like MGSM and MMLU-ProX, which test performance across multiple languages to identify models with the strongest cross-lingual capabilities.
What do MGSM and MMLU-ProX evaluate?
MGSM evaluates multilingual math reasoning, while MMLU-ProX is a broader multilingual professional benchmark that captures cross-language knowledge and reasoning beyond translated arithmetic.
How do multilingual benchmarks differ from English-only benchmarks?
Multilingual benchmarks test model performance across many languages simultaneously, revealing how well capabilities transfer beyond English, where most training data is concentrated.
Multilingual benchmark updates
Which model handles your language best? Updated weekly.
One email each week. Unsubscribe anytime.