Skip to main content
Multilingual benchmark report

Best LLMs for MultilingualJuly 2026 Leaderboard

As of July 2026, the top multilingual model on the BenchLM leaderboard is Qwen3.7 Max with a weighted multilingual score of 100.

Data refreshed:

Performance across multiple languages

Decision lens: use provisional-ranked mode for broader public evidence and verified-ranked mode for source-only comparisons. A model can move between views as evidence coverage changes.

Data refreshed
July 21, 2026
Provisional-ranked
13 of 290 models
Verified-ranked
12 of 290 models
Weighted evidence
1 of 2 benchmarks
2 tracked benchmarks

MGSM, MMLU-ProX

Best Multilingual picks

BenchLM summaries for multilingual plus the practical tradeoffs users check next: open weights, price, speed, latency, and context.

How BenchLM scores these

Multilingual Leaderboard

Primary score: weighted multilingual score. Higher values rank first. Use the Show metric control to change the value shown in each row.

Updated Embed leaderboard

Switch between provisional-ranked and verified-ranked modes to compare the broader public dataset with sourced-only rankings.

Filters
Provisional-ranked mode includes source-unverified non-generated benchmark evidence.P = provisional benchmark row
Rank / modelWeighted Multilingual
1
Qwen3.7 MaxAlibaba · Closed
100%
2
Claude Opus 4.5Anthropic · Closed
82.9%
3
Qwen3.7 PlusAlibaba · Closed
78.9%
4
Qwen3.6 PlusAlibaba · Closed
69.7%
5
Qwen3.5 397BAlibaba · Open weight
69.7%
6
GLM-5Z.AI · Open weight
48.7%
7
Nemotron 3 UltraNVIDIA · Open weight
47.4%
8
Kimi K2.5Moonshot AI · Open weight
38.2%
9
Qwen3.5-122B-A10BAlibaba · Open weight
36.8%
10
Qwen3.5-27BAlibaba · Open weight
36.8%
11
Qwen3.5-35B-A3BAlibaba · Open weight
21.1%
12
Qwen3 235B 2507Alibaba · Open weight
1%
13
Claude Mythos 5Anthropic · Closed
0%

Top AI Models for MultilingualJuly 2026

As of July 2026, Qwen3.7 Max leads the provisional multilingual leaderboard with a score of 100.0%, followed by Claude Opus 4.5 (82.9%) and Qwen3.7 Plus (78.9%). BenchLM is currently showing 13 provisional-ranked models and 12 verified-ranked models in this category.

What changed

Claude Mythos Preview leads multilingual with the most consistent cross-language scores.

GPT-5.4 close second, strong on MMLU-ProX across all tested languages.

Claude Opus 4.6 holds #3, with particularly strong MGSM performance.

Top models by benchmark

Broad multilingual professional benchmark across many languages(100% of category score)

RankModelReported score

Score in Context

What these scores mean

Multilingual carries a 7% weight in overall scoring. The weighted score blends MGSM (multilingual math reasoning) and MMLU-ProX (cross-language professional knowledge). This category reveals how well model capabilities transfer beyond English, where most training data is concentrated.

Known limitations

Only two benchmarks cover this category, which limits the signal. MGSM tests math reasoning specifically, not general language quality. Languages tested are limited — low-resource languages remain untested. A model scoring well here may still struggle with less common languages or dialects.

How we weight

Multilingual carries a 7% weight in BenchLM.ai's overall scoring. Cross-language performance reveals how well model capabilities transfer beyond English. See the multilingual leaderboard or compare with knowledge benchmarks.

Leaderboards exclude benchmark rows that BenchLM generated from other scores or cloned from reference models. When a weighted benchmark is missing after that filter, the category falls back to the remaining trustworthy public rows instead of filling the gap with synthetic values.

The full scoring rules, freshness handling, and runtime/pricing caveats live on the BenchLM methodology page.

Scroll horizontally to read the full evidence ledger.

Multilingual benchmark weights, ranking status, and descriptions
BenchmarkWeightStatusDescription
MGSMDisplay onlyGrade school math problems translated into 10 diverse languages plus English
MMLU-ProX100%WeightedBroad multilingual professional benchmark across many languages

About Multilingual Benchmarks

Grade school math problems translated into 10 diverse languages plus English

Common questions

What is the best LLM for multilingual tasks?

The top multilingual LLMs are ranked by benchmarks like MGSM and MMLU-ProX, which test performance across multiple languages to identify models with the strongest cross-lingual capabilities.

What do MGSM and MMLU-ProX evaluate?

MGSM evaluates multilingual math reasoning, while MMLU-ProX is a broader multilingual professional benchmark that captures cross-language knowledge and reasoning beyond translated arithmetic.

How do multilingual benchmarks differ from English-only benchmarks?

Multilingual benchmarks test model performance across many languages simultaneously, revealing how well capabilities transfer beyond English, where most training data is concentrated.

Multilingual benchmark updates

Which model handles your language best? Updated weekly.

One email each week. Unsubscribe anytime.

Related