Skip to main content
Math benchmark report

Best LLMs for MathJuly 2026 Leaderboard

As of July 2026, the top math model on the BenchLM leaderboard is Kimi K2.6 with a weighted math score of 72.3.

Data refreshed:

Mathematical reasoning and problem solving

Decision lens: use provisional-ranked mode for broader public evidence and verified-ranked mode for source-only comparisons. A model can move between views as evidence coverage changes.

Data refreshed
July 21, 2026
Provisional-ranked
7 of 290 models
Verified-ranked
7 of 290 models
Weighted evidence
0 of 9 benchmarks
9 tracked benchmarks

AIME 2023, AIME 2024, AIME 2025, AIME25 (Arcee), HMMT Feb 2023, HMMT Feb 2024, HMMT Feb 2025, BRUMO 2025, MATH-500

Best Math picks

BenchLM summaries for math plus the practical tradeoffs users check next: open weights, price, speed, latency, and context.

How BenchLM scores these

Math Leaderboard

Primary score: weighted math score. Higher values rank first. Use the Show metric control to change the value shown in each row.

Updated Embed leaderboard

Switch between provisional-ranked and verified-ranked modes to compare the broader public dataset with sourced-only rankings.

Filters
Provisional-ranked mode includes source-unverified non-generated benchmark evidence.P = provisional benchmark row
Rank / modelWeighted Math
1
Kimi K2.6Moonshot AI · Open weight
72.3%
2
Claude Opus 4.8Anthropic · Closed
66.9%
3
GLM-5.1Z.AI · Open weight
64.6%
4
Qwen3.6 PlusAlibaba · Closed
63%
5
Kimi K2.5Moonshot AI · Open weight
62.8%
6
Claude Opus 4.5Anthropic · Closed
58.6%
7
GLM-5Z.AI · Open weight
57.2%

Top AI Models for MathJuly 2026

As of July 2026, Kimi K2.6 leads the provisional math leaderboard with a score of 72.3%, followed by Claude Opus 4.8 (66.9%) and GLM-5.1 (64.6%). BenchLM is currently showing 7 provisional-ranked models and 7 verified-ranked models in this category.

What changed

Claude Mythos Preview leads math with top BRUMO and MATH-500 scores.

GPT-5.4 close second, with near-perfect AIME scores.

Gemini 3.1 Pro strong third — best value option for math-heavy workloads.

Top models by benchmark

RankModelReported score
1GLM-5.299.2
2Inkling97.1
4GLM-595.8

Score in Context

What these scores mean

Math carries a 5% weight in overall scoring — relatively low because frontier models have saturated the main competition benchmarks. AIME and HMMT scores are 95-99% across top models. The weighted score now relies on BRUMO and MATH-500, which still show meaningful separation.

Known limitations

AIME and HMMT are effectively solved by AI — they are displayed for reference but no longer factor into the weighted score. If math reasoning is critical for your use case, look at BRUMO scores specifically, and consider models with explicit reasoning capabilities (chain-of-thought). See the AIME & HMMT explainer.

How we weight

Mathematics carries a 5% weight in BenchLM.ai's overall scoring. Frontier models score 95-99% on AIME and HMMT — competition math is effectively solved by AI.

AIME and HMMT are still displayed for reference but no longer factor into the weighted score due to saturation. BRUMO and MATH-500 show more meaningful separation. If mathematical reasoning is critical, prioritize models with explicit reasoning capabilities. See the math leaderboard or read the AIME & HMMT explainer.

Leaderboards exclude benchmark rows that BenchLM generated from other scores or cloned from reference models. When a weighted benchmark is missing after that filter, the category falls back to the remaining trustworthy public rows instead of filling the gap with synthetic values.

The full scoring rules, freshness handling, and runtime/pricing caveats live on the BenchLM methodology page.

Scroll horizontally to read the full evidence ledger.

Math benchmark weights, ranking status, and descriptions
BenchmarkWeightStatusDescription
AIME 2023Display onlyHigh school mathematics competition
AIME 2024Display onlyHigh school mathematics competition
AIME 2025Display onlyHigh school mathematics competition
AIME25 (Arcee)Display onlyDisplay-only AIME25 reference from Arcee AI's Trinity-Large-Thinking launch chart.
HMMT Feb 2023Display onlyCollegiate mathematics competition
HMMT Feb 2024Display onlyCollegiate mathematics competition
HMMT Feb 2025Display onlyCollegiate mathematics competition
BRUMO 2025Display onlyUniversity-level mathematics olympiad
MATH-500Display onlyCurated 500-problem subset of the MATH dataset covering algebra, geometry, number theory, and more

About Math Benchmarks

High school mathematics competition

Common questions

What is the best LLM for math?

The best LLMs for math are ranked by competition-level benchmarks like AIME and HMMT, with top models achieving strong scores on problems from real math competitions.

How are math benchmarks scored for LLMs?

Math benchmarks score models on correctness of final answers to problems ranging from algebra to advanced competition mathematics, testing both calculation and mathematical reasoning.

What benchmarks test math ability in AI models?

Major math benchmarks include AIME (competition-level problems), HMMT (Harvard-MIT tournament problems), and BRUMO, each testing progressively harder mathematical reasoning.

Math benchmark updates

Math model rankings change weekly. Stay current.

One email each week. Unsubscribe anytime.

Related