# Long-context translation evaluation (Cohere) (Long-context translation)

> Cohere's long-document translation evaluation: two book chapters translated in a single call, scored per paragraph with xCOMET-XL.

Canonical page: https://benchlm.ai/benchmarks/long-context-translation

- Category: [Multilingual](/multilingual)
- Last updated: September 27, 2026

## About Long-context translation

- Year: 2026
- Tasks: Two book chapters per call
- Format: xCOMET-XL quality score
- Difficulty: Long-document translation
- Paper: [Introducing North Small Translate](https://cohere.com/blog/north-small-translate)

A provider-designed evaluation published with the North Small Translate launch. Quality is measured for each paragraph in isolation via xCOMET-XL after translating two chapters of a book in one call. BenchLM keeps it display-only because the source texts and scoring pipeline are not public.

Long-context translation is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (1 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [North Small Translate](/models/north-small-translate-1-0) | Cohere | 48.9% |

## FAQ

### What does Long-context translation measure?

Cohere's long-document translation evaluation: two book chapters translated in a single call, scored per paragraph with xCOMET-XL.

### Which model scores highest on Long-context translation?

North Small Translate by Cohere currently leads with a score of 48.9% on Long-context translation.

### How many models are evaluated on Long-context translation?

1 AI models have been evaluated on Long-context translation on BenchLM.

### Does Long-context translation affect BenchLM's overall score?

Not directly. Long-context translation is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.
