Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes

MATH-500 Problem Set (MATH-500)

A curated subset of 500 problems from the MATH dataset, covering algebra, counting and probability, geometry, intermediate algebra, number theory, prealgebra, and precalculus.

Data verified 31 confirmed releases in the last 30 daysSee provider release alerts

Benchmark score on MATH-500 — September 10, 2026

We mirror the published score view for MATH-500. LongCat-Flash-Lite-Sparse leads the public snapshot at 95.8%, followed by MiniCPM5-2B (94.6%) and MiniCPM5-1B (91.6%). We do not use these results to rank models overall.

6 modelsMathStaleDisplay onlyUpdated September 10, 2026

Benchmark score table (6 models)

Score
1
LongCat-Flash-Lite-SparseMeituan · Open weight
95.8%
2
MiniCPM5-2BOpenBMB · Open weight
94.6%
3
MiniCPM5-1BOpenBMB · Open weight
91.6%
4
LFM2.5-8B-A1BLiquidAI · Open weight
88.8%
5
Kanana-2 1.3B InstructKakao · Open weight
61.4%
6
Kanana-2 3B InstructKakao · Open weight
61.2%

The published MATH-500 snapshot places LongCat-Flash-Lite-Sparse first at 95.8%. The third row is 4.2 points behind. The broader top-10 range is 34.6 points, so the table still separates the published systems.

6 models have been evaluated on MATH-500. The benchmark falls in the Math category. This category carries a 5% weight in BenchLM.ai's overall scoring system. MATH-500 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About MATH-500

Year

2021

Tasks

500 problems

Format

Free-form mathematical answers

Difficulty

High school to undergraduate

MATH-500 is one of the most widely cited math benchmarks. It is nearing saturation with top reasoning models scoring 96-99%, making it less useful for differentiating frontier models but still a standard baseline.

BenchLM freshness & provenance

Version

MATH-500 2021

Refresh cadence

Static

Staleness state

Stale

Question availability

Public benchmark set

StaleDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does MATH-500 measure?

A curated subset of 500 problems from the MATH dataset, covering algebra, counting and probability, geometry, intermediate algebra, number theory, prealgebra, and precalculus.

Which model scores highest on MATH-500?

LongCat-Flash-Lite-Sparse by Meituan currently leads with a score of 95.8% on MATH-500.

How many models are evaluated on MATH-500?

6 AI models have been evaluated on MATH-500 on BenchLM.

Last updated: September 10, 2026 · BenchLM version MATH-500 2021

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.