MATH-500 Problem Set (MATH-500)
A curated subset of 500 problems from the MATH dataset, covering algebra, counting and probability, geometry, intermediate algebra, number theory, prealgebra, and precalculus.
Data verified 31 confirmed releases in the last 30 daysSee provider release alertsBenchmark score on MATH-500 — September 10, 2026
We mirror the published score view for MATH-500. LongCat-Flash-Lite-Sparse leads the public snapshot at 95.8%, followed by MiniCPM5-2B (94.6%) and MiniCPM5-1B (91.6%). We do not use these results to rank models overall.
LongCat-Flash-Lite-Sparse
Meituan
MiniCPM5-2B
OpenBMB
MiniCPM5-1B
OpenBMB
6 modelsMathStaleDisplay onlyUpdated September 10, 2026
Benchmark score table (6 models)
ScoreThe published MATH-500 snapshot places LongCat-Flash-Lite-Sparse first at 95.8%. The third row is 4.2 points behind. The broader top-10 range is 34.6 points, so the table still separates the published systems.
6 models have been evaluated on MATH-500. The benchmark falls in the Math category. This category carries a 5% weight in BenchLM.ai's overall scoring system. MATH-500 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About MATH-500
Year
2021
Tasks
500 problems
Format
Free-form mathematical answers
Difficulty
High school to undergraduate
MATH-500 is one of the most widely cited math benchmarks. It is nearing saturation with top reasoning models scoring 96-99%, making it less useful for differentiating frontier models but still a standard baseline.
BenchLM freshness & provenance
Version
MATH-500 2021
Refresh cadence
Static
Staleness state
Stale
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does MATH-500 measure?
A curated subset of 500 problems from the MATH dataset, covering algebra, counting and probability, geometry, intermediate algebra, number theory, prealgebra, and precalculus.
Which model scores highest on MATH-500?
LongCat-Flash-Lite-Sparse by Meituan currently leads with a score of 95.8% on MATH-500.
How many models are evaluated on MATH-500?
6 AI models have been evaluated on MATH-500 on BenchLM.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.