Benchmark profile
MATH-500 Problem Set (MATH-500)
A curated subset of 500 problems from the MATH dataset, covering algebra, counting and probability, geometry, intermediate algebra, number theory, prealgebra, and precalculus.
Data verifiedBenchmark score on MATH-500 — July 20, 2026
BenchLM mirrors the published score view for MATH-500. MiniCPM5-1B leads the public snapshot at 91.6% , followed by LFM2.5-8B-A1B (88.8%). BenchLM does not use these results to rank models overall.
MiniCPM5-1B
OpenBMB
minicpm5-1b
LFM2.5-8B-A1B
LiquidAI
lfm2-5-8b-a1b
Benchmark score table (2 models)
ScoreAbout MATH-500
Year
2021
Tasks
500 problems
Format
Free-form mathematical answers
Difficulty
High school to undergraduate
MATH-500 is one of the most widely cited math benchmarks. It is nearing saturation with top reasoning models scoring 96-99%, making it less useful for differentiating frontier models but still a standard baseline.
BenchLM freshness & provenance
Version
MATH-500 2021
Refresh cadence
Static
Staleness state
Stale
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does MATH-500 measure?
A curated subset of 500 problems from the MATH dataset, covering algebra, counting and probability, geometry, intermediate algebra, number theory, prealgebra, and precalculus.
Which model scores highest on MATH-500?
MiniCPM5-1B by OpenBMB currently leads with a score of 91.6% on MATH-500.
How many models are evaluated on MATH-500?
2 AI models have been evaluated on MATH-500 on BenchLM.
Compare Top Models on MATH-500
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.