# Best LLMs for Multilingual — September 2026 Leaderboard

> As of September 2026, Qwen3.7 Max leads BenchLM's multilingual leaderboard with a weighted score of 100.

- **Last verified:** September 4, 2026
- Canonical page: https://benchlm.ai/multilingual
- **Ranking coverage:** 13 category-ranked models from 411 tracked models
- **Category weight:** 7% of the overall BenchLM score

## Current ranking

| Rank | Model | Creator | Weighted score | Published category rows | Exact-source rows (all categories) |
|------|-------|---------|----------------|----------------|-------------------|
| 1 | [Qwen3.7 Max](/models/qwen3-7-max) | Alibaba | 100 | 5 | 47 total |
| 2 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 82.9 | 2 | 50 total |
| 3 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 78.9 | 5 | 61 total |
| 4 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 69.7 | 2 | 53 total |
| 5 | [Qwen3.5 397B](/models/qwen3-5-397b) | Alibaba | 69.7 | 2 | 35 total |
| 6 | [GLM-5](/models/glm-5) | Z.AI | 48.7 | 2 | 41 total |
| 7 | [Nemotron 3 Ultra](/models/nemotron-3-ultra) | NVIDIA | 47.4 | 1 | 26 total |
| 8 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 38.2 | 2 | 49 total |
| 9 | [Qwen3.5-27B](/models/qwen3-5-27b) | Alibaba | 36.8 | 1 | 21 total |
| 10 | [Qwen3.5-122B-A10B](/models/qwen3-5-122b-a10b) | Alibaba | 36.8 | 1 | 22 total |
| 11 | [Qwen3.5-35B-A3B](/models/qwen3-5-35b-a3b) | Alibaba | 21.1 | 1 | 21 total |
| 12 | [Qwen3 235B 2507](/models/qwen3-235b-2507) | Alibaba | 1 | 1 | 4 total |
| 13 | [Claude Mythos 5](/models/claude-mythos-5) | Anthropic | 0 | 1 | 14 total |

## Decision-ready shortlist

- #1 [Qwen3.7 Max](/models/qwen3-7-max) — 100 weighted score, Proprietary, 1M context.
- #2 [Claude Opus 4.5](/models/claude-opus-4-5) — 82.9 weighted score, Proprietary, 200K context.
- #3 [Qwen3.7 Plus](/models/qwen3-7-plus) — 78.9 weighted score, Proprietary, 1M context.
- #4 [Qwen3.6 Plus](/models/qwen3-6-plus) — 69.7 weighted score, Proprietary, 1M context.
- #5 [Qwen3.5 397B](/models/qwen3-5-397b) — 69.7 weighted score, Open Weight, 128K context.

## Benchmarks in this category

### [MGSM](/benchmarks/mgsm) (Multilingual Grade School Math)

A multilingual benchmark that translates 250 grade school math problems from GSM8K into 10 typologically diverse languages: Bengali, German, Spanish, French, Japanese, Russian, Swahili, Telugu, Thai, and Chinese.

- Ranking status: Display only
- Year: 2022
- Format: Math word problems
- Difficulty: Grade school math, multilingual

### [MMLU-ProX](/benchmarks/mmluprox) (MMLU-ProX)

A multilingual extension of professional-level academic evaluation across many languages.

- Ranking status: Weighted (100% of this category)
- Year: 2025
- Format: Multilingual multiple choice
- Difficulty: Professional multilingual
