Skip to main content
BenchLM

Multilingual Grade School Math (MGSM)

We show this table for reference; we do not rank on it.

Data verified 34 confirmed releases in the last 30 daysFollow model changes

A multilingual benchmark that translates 250 grade school math problems from GSM8K into 10 typologically diverse languages: Bengali, German, Spanish, French, Japanese, Russian, Swahili, Telugu, Thai, and Chinese.

Benchmark score on MGSM — September 27, 2026

We compile the MGSM rows from provider self-reports. DeepSeek V4 Flash Base leads the table at 85.7%, followed by DeepSeek V4 Pro Base (84.4%). We do not use these results to rank models overall.

1Open weights

DeepSeek V4 Flash Base

DeepSeek

85.7%
Context 1M
2Open weights

DeepSeek V4 Pro Base

DeepSeek

84.4%
Context 1M

2 modelsMultilingualStaleDisplay onlyUpdated September 27, 2026

Benchmark score table (2 models)

Score
1
DeepSeek V4 Flash BaseDeepSeek · Open weight
85.7%
2
DeepSeek V4 Pro BaseDeepSeek · Open weight
84.4%

About MGSM

Year

2022

Tasks

250 problems × 11 languages

Format

Math word problems

Difficulty

Grade school math, multilingual

MGSM evaluates mathematical reasoning across languages, revealing that performance can vary significantly across languages, with lower-resource languages (Bengali, Swahili, Telugu) typically showing the largest gaps.

Freshness and provenance

Version

MGSM 2022

Refresh cadence

Static

Staleness state

Stale

Question availability

Public benchmark set

StaleDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does MGSM measure?

A multilingual benchmark that translates 250 grade school math problems from GSM8K into 10 typologically diverse languages: Bengali, German, Spanish, French, Japanese, Russian, Swahili, Telugu, Thai, and Chinese.

Which model scores highest on MGSM?

DeepSeek V4 Flash Base by DeepSeek currently leads with a score of 85.7% on MGSM.

How many models are evaluated on MGSM?

2 AI models have been evaluated on MGSM on BenchLM.

Last updated: September 27, 2026 · BenchLM version MGSM 2022

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.