# MATH-500 Problem Set (MATH-500)

> A curated subset of 500 problems from the MATH dataset, covering algebra, counting and probability, geometry, intermediate algebra, number theory, prealgebra, and precalculus.

Canonical page: https://benchlm.ai/benchmarks/math-500

- Category: [Mathematics](/math)
- Last updated: September 10, 2026

## About MATH-500

- Year: 2021
- Tasks: 500 problems
- Format: Free-form mathematical answers
- Difficulty: High school to undergraduate
- Paper: [Measuring Mathematical Problem Solving With the MATH Dataset](https://arxiv.org/abs/2103.03874)

MATH-500 is one of the most widely cited math benchmarks. It is nearing saturation with top reasoning models scoring 96-99%, making it less useful for differentiating frontier models but still a standard baseline.

MATH-500 is currently displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Leaderboard (6 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [LongCat-Flash-Lite-Sparse](/models/longcat-flash-lite-sparse) | Meituan | 95.8% |
| 2 | [MiniCPM5-2B](/models/minicpm5-2b) | OpenBMB | 94.6% |
| 3 | [MiniCPM5-1B](/models/minicpm5-1b) | OpenBMB | 91.6% |
| 4 | [LFM2.5-8B-A1B](/models/lfm2-5-8b-a1b) | LiquidAI | 88.8% |
| 5 | [Kanana-2 1.3B Instruct](/models/kanana-2-1-3b-instruct) | Kakao | 61.4% |
| 6 | [Kanana-2 3B Instruct](/models/kanana-2-3b-instruct) | Kakao | 61.2% |

## FAQ

### What does MATH-500 measure?

A curated subset of 500 problems from the MATH dataset, covering algebra, counting and probability, geometry, intermediate algebra, number theory, prealgebra, and precalculus.

### Which model scores highest on MATH-500?

LongCat-Flash-Lite-Sparse by Meituan currently leads with a score of 95.8% on MATH-500.

### How many models are evaluated on MATH-500?

6 AI models have been evaluated on MATH-500 on BenchLM.

### Does MATH-500 affect BenchLM's overall score?

Not directly. MATH-500 is still displayed on BenchLM for reference, but it is excluded from the weighted scoring formula.

## Compare Top Models on MATH-500

- [LongCat-Flash-Lite-Sparse vs MiniCPM5-2B](/compare/longcat-flash-lite-sparse-vs-minicpm5-2b)
- [MiniCPM5-2B vs MiniCPM5-1B](/compare/minicpm5-1b-vs-minicpm5-2b)
- [MiniCPM5-1B vs LFM2.5-8B-A1B](/compare/lfm2-5-8b-a1b-vs-minicpm5-1b)
- [LFM2.5-8B-A1B vs Kanana-2 1.3B Instruct](/compare/kanana-2-1-3b-instruct-vs-lfm2-5-8b-a1b)
