# PolyMath

> A multilingual mathematical reasoning benchmark that tests whether math performance transfers across languages rather than only in English.

Canonical page: https://benchlm.ai/benchmarks/polymath

- Category: [Multilingual](/multilingual)
- Last updated: September 27, 2026

## About PolyMath

- Year: 2026
- Tasks: Multilingual math problems
- Format: Cross-lingual mathematical reasoning
- Difficulty: Advanced multilingual reasoning
- Paper: [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6)

PolyMath isolates cross-lingual math transfer rather than general chat quality. It is useful for spotting models that keep surface fluency in other languages but lose structured reasoning quality.

PolyMath is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (3 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Qwen3.7 Max](/models/qwen3-7-max) | Alibaba | 86.5% |
| 2 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 84.0% |
| 3 | [K-EXAONE 2.0](/models/k-exaone-2-0) | LG AI Research | 71.3% |

## FAQ

### What does PolyMath measure?

A multilingual mathematical reasoning benchmark that tests whether math performance transfers across languages rather than only in English.

### Which model scores highest on PolyMath?

Qwen3.7 Max by Alibaba currently leads with a score of 86.5% on PolyMath.

### How many models are evaluated on PolyMath?

3 AI models have been evaluated on PolyMath on BenchLM.

### Does PolyMath affect BenchLM's overall score?

Not directly. PolyMath is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on PolyMath

- [Qwen3.7 Max vs Qwen3.7 Plus](/compare/qwen3-7-max-vs-qwen3-7-plus)
- [Qwen3.7 Plus vs K-EXAONE 2.0](/compare/k-exaone-2-0-vs-qwen3-7-plus)
