Skip to main content
BenchLM

PolyMath

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

A multilingual mathematical reasoning benchmark that tests whether math performance transfers across languages rather than only in English.

Benchmark score on PolyMath — September 27, 2026

We compile the PolyMath rows from provider self-reports. Qwen3.7 Max leads the table at 86.5%, followed by Qwen3.7 Plus (84.0%) and K-EXAONE 2.0 (71.3%). We do not use these results to rank models overall.

3 modelsMultilingualCurrentDisplay onlyUpdated September 27, 2026

Benchmark score table (3 models)

Score
1
Qwen3.7 MaxAlibaba · Closed
86.5%
2
Qwen3.7 PlusAlibaba · Closed
84.0%
3
K-EXAONE 2.0LG AI Research · Open weight
71.3%

Among the reported PolyMath rows, Qwen3.7 Max is first at 86.5%. The third row is 15.2 points behind. The broader top-10 range is 15.2 points, so the table still separates the published systems.

3 models have been evaluated on PolyMath. The benchmark falls in the Multilingual category. PolyMath is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About PolyMath

Year

2026

Tasks

Multilingual math problems

Format

Cross-lingual mathematical reasoning

Difficulty

Advanced multilingual reasoning

PolyMath isolates cross-lingual math transfer rather than general chat quality. It is useful for spotting models that keep surface fluency in other languages but lose structured reasoning quality.

Freshness and provenance

Version

PolyMath 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does PolyMath measure?

A multilingual mathematical reasoning benchmark that tests whether math performance transfers across languages rather than only in English.

Which model scores highest on PolyMath?

Qwen3.7 Max by Alibaba currently leads with a score of 86.5% on PolyMath.

How many models are evaluated on PolyMath?

3 AI models have been evaluated on PolyMath on BenchLM.

Last updated: September 27, 2026 · BenchLM version PolyMath 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.