# Grade School Math 8K (GSM8K)

> A grade-school mathematical reasoning benchmark reported in DeepSeek-V4 base-model evaluations.

Canonical page: https://benchlm.ai/benchmarks/gsm8k

- Category: [Mathematics](/math)
- Last updated: September 15, 2026

## About GSM8K

- Year: 2026
- Tasks: Grade-school math word problems
- Format: Exact match
- Difficulty: Grade-school math
- Paper: [DeepSeek-V4 Technical Report](https://huggingface.co/deepseek-ai/DeepSeek-V4-Pro/blob/main/DeepSeek_V4.pdf)

BenchLM stores GSM8K as a display-only provider-table row when exact values are published in DeepSeek-V4 evaluations.

GSM8K is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (2 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [Celeris-1](/models/celeris-1) | Celeris | 93.7% |
| 2 | [Soofi S 30B-A3B](/models/soofi-s-30b-a3b) | Soofi Project | 86.1% |

## FAQ

### What does GSM8K measure?

A grade-school mathematical reasoning benchmark reported in DeepSeek-V4 base-model evaluations.

### Which model scores highest on GSM8K?

Celeris-1 by Celeris currently leads with a score of 93.7% on GSM8K.

### How many models are evaluated on GSM8K?

2 AI models have been evaluated on GSM8K on BenchLM.

### Does GSM8K affect BenchLM's overall score?

Not directly. GSM8K is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on GSM8K

- [Celeris-1 vs Soofi S 30B-A3B](/compare/celeris-1-vs-soofi-s-30b-a3b)
