# WMT26 General Machine Translation (all languages) (WMT26)

> Machine-translation quality on the WMT26 general translation task, averaged across all evaluated languages and scored 0-100 by an LLM judge.

Canonical page: https://benchlm.ai/benchmarks/wmt26

- Category: [Multilingual](/multilingual)
- Last updated: September 15, 2026

## About WMT26

- Year: 2026
- Tasks: General-domain translation across 50+ languages (32 high-resource plus 18 additional)
- Format: 0-100 judged translation quality, averaged across languages
- Difficulty: Professional machine translation
- Paper: [WMT26 General Machine Translation Task](https://www2.statmt.org/wmt26/translation-task.html)

BenchLM currently stores provider-run WMT26 evaluations. Cohere's September 2026 run covers 50+ languages with GPT-5.6 Sol as the judge on a 0-100 scale, where 80-100 means perfect or minor errors. Because the judge, language set, and prompts are provider-chosen, rows stay display-only and are not directly comparable to official WMT human evaluations.

WMT26 is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (1 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [North Small Translate](/models/north-small-translate-1-0) | Cohere | 83.6% |

## FAQ

### What does WMT26 measure?

Machine-translation quality on the WMT26 general translation task, averaged across all evaluated languages and scored 0-100 by an LLM judge.

### Which model scores highest on WMT26?

North Small Translate by Cohere currently leads with a score of 83.6% on WMT26.

### How many models are evaluated on WMT26?

1 AI models have been evaluated on WMT26 on BenchLM.

### Does WMT26 affect BenchLM's overall score?

Not directly. WMT26 is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
