Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes
BenchLM recommendation

Best LLMs for Translation in 2026

Data verified

As of September 10, 2026, the top model in best llms for translation on the BenchLM leaderboard is Qwen3.7 Max with a score of 100.

Bottom line: Claude Fable 5, Claude Mythos 5, and Gemini 3.1 Pro are tied at the top of the multilingual category — pick by price and ecosystem for major-language translation, and test low-resource pairs yourself.

About this ranking

Last verified: September 10, 2026

Translation quality tracks the multilingual category: benchmarks that test comprehension and generation across languages. The frontier models at the top of this table are effectively tied on major-language pairs; the gaps show up in lower-resource languages, idiom, and domain terminology.

Unless noted otherwise, ranking surfaces on this page use BenchLM’s provisional leaderboard lane rather than the stricter sourced-only verified leaderboard.

Qwen3.7 Max leads this ranking with a score of 100, followed by Claude Opus 4.5 (82.9) and Qwen3.7 Plus (78.9). There is a significant gap between the leading models and the rest of the field.

The best open-weight option is Qwen3.5 397B (ranked #5 with a score of 69.7). While proprietary models lead, open-weight options are within striking distance for teams willing to trade a few points of performance for full model control.

This ranking uses provisional weighted averages across the scoring benchmarks in multilingual. For detailed model profiles, click any model name below. To compare two specific models head-to-head, use the "vs #" links.

What changed

Qwen3.7 Max leads the live llms for translation ranking at 100 with Estimated evidence.

Claude Opus 4.5 ranks #2 at 82.9 with Estimated evidence.

Qwen3.7 Plus ranks #3 at 78.9 with Estimated evidence.

How to choose

Full Rankings (12 models)

1
Qwen3.7 Max
Alibaba·Proprietary·1M

100

prov. avg

2
Claude Opus 4.5
Anthropic·Proprietary·200K

82.9

prov. avg

3
Qwen3.7 Plus
Alibaba·Proprietary·1M

78.9

prov. avg

4
Qwen3.6 Plus
Alibaba·Proprietary·1M

69.7

prov. avg

5
Qwen3.5 397B
Alibaba·Open Weight·128K

69.7

prov. avg

6
GLM-5
Z.AI·Open Weight·200K

48.7

prov. avg

7
Nemotron 3 Ultra
NVIDIA·Open Weight·1M

47.4

prov. avg

8
Kimi K2.5
Moonshot AI·Open Weight·256K

38.2

prov. avg

9
Qwen3.5-27B
Alibaba·Open Weight·262K

36.8

prov. avg

10
Qwen3.5-122B-A10B
Alibaba·Open Weight·262K

36.8

prov. avg

11
Qwen3.5-35B-A3B
Alibaba·Open Weight·262K

21.1

prov. avg

12
Qwen3 235B 2507
Alibaba·Open Weight·128K

1

prov. avg

Key Takeaways

The top model is Qwen3.7 Max by Alibaba with a provisional score of 100.

The best open-weight model is Qwen3.5 397B at position #5.

12 models are included in this ranking.

Score in Context

What these scores mean

The multilingual score blends cross-language comprehension and generation benchmarks like MGSM and MMLU-ProX. It is the closest measured proxy for translation strength BenchLM tracks.

Known limitations

Benchmarks over-represent high-resource languages. For low-resource pairs, dialects, or domain terminology (legal, medical), run your own evaluation set — leaderboard gaps do not transfer reliably.

Best LLMs for Translation FAQ

What is the best LLM for translation?

Claude Fable 5, Claude Mythos 5, and Gemini 3.1 Pro are tied at the top of BenchLM's multilingual category. For most translation work Gemini 3.1 Pro is the practical pick — leader-tier quality at $2/$12 per million tokens, roughly a fifth of Fable 5's price.

Are LLMs better than Google Translate?

For context-heavy translation — documents, marketing copy, anything where tone matters — frontier LLMs generally produce more natural output because they use surrounding context and follow style instructions. For quick single sentences, dedicated translation tools remain faster and cheaper.

What is the best open-source model for translation?

Alibaba's Qwen rows are the strongest open-weight multilingual performers BenchLM tracks, which fits their heavily multilingual training focus. See the Qwen rankings and the open-source leaderboard for current scores per model.

How should I evaluate translation quality myself?

Build a 30-50 segment test set from your real content across your target pairs, translate with 2-3 shortlisted models, and have a native speaker rank blind. Benchmark scores shortlist correctly, but domain terminology and tone preferences are yours to verify.

Last updated: September 10, 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.