Best LLMs for Math — July 2026 Leaderboard
As of July 2026, the top math model on the BenchLM leaderboard is Kimi K2.6 with a weighted math score of 72.3.
Data refreshed:
Mathematical reasoning and problem solving
Decision lens: use provisional-ranked mode for broader public evidence and verified-ranked mode for source-only comparisons. A model can move between views as evidence coverage changes.
- Data refreshed
- July 21, 2026
- Provisional-ranked
- 7 of 290 models
- Verified-ranked
- 7 of 290 models
- Weighted evidence
- 0 of 9 benchmarks
9 tracked benchmarks
AIME 2023, AIME 2024, AIME 2025, AIME25 (Arcee), HMMT Feb 2023, HMMT Feb 2024, HMMT Feb 2025, BRUMO 2025, MATH-500
Evidence set: AIME 2023, AIME 2024, AIME 2025, AIME25 (Arcee), HMMT Feb 2023, HMMT Feb 2024, HMMT Feb 2025, BRUMO 2025, MATH-500
Best Math picks
BenchLM summaries for math plus the practical tradeoffs users check next: open weights, price, speed, latency, and context.
Math Leaderboard
Primary score: weighted math score. Higher values rank first. Use the Show metric control to change the value shown in each row.
Switch between provisional-ranked and verified-ranked modes to compare the broader public dataset with sourced-only rankings.
Filters
| 72.3% | 65 | — | — | — | — | — | — | — | — | — | |
2 Claude Opus 4.8 Anthropic | 66.9% | 82 | — | — | — | — | — | — | — | — | — |
| 64.6% | 64 | — | — | 93.3%P | — | — | — | — | 87%P | 97.4%P | |
4 Qwen3.6 Plus Alibaba | 63% | 64 | — | — | — | — | — | — | — | — | — |
| 62.8% | 56 | 77%P | 79%P | 96.1% | 96.3% | 73%P | 75%P | 74%P | 76%P | 82%P | |
6 Claude Opus 4.5 Anthropic | 58.6% | 60 | 99%P | 99%P | 98%P | — | 95%P | 97%P | 96%P | 96%P | 89%P |
7 | 57.2% | 64 | 88%P | 90%P | 93.3%P | 93.3% | 84%P | 86%P | 85%P | 87%P | 97.4%P |
Top AI Models for Math — July 2026
As of July 2026, Kimi K2.6 leads the provisional math leaderboard with a score of 72.3%, followed by Claude Opus 4.8 (66.9%) and GLM-5.1 (64.6%). BenchLM is currently showing 7 provisional-ranked models and 7 verified-ranked models in this category.
What changed
Claude Mythos Preview leads math with top BRUMO and MATH-500 scores.
GPT-5.4 close second, with near-perfect AIME scores.
Gemini 3.1 Pro strong third — best value option for math-heavy workloads.
Top models by benchmark
Score in Context
What these scores mean
Math carries a 5% weight in overall scoring — relatively low because frontier models have saturated the main competition benchmarks. AIME and HMMT scores are 95-99% across top models. The weighted score now relies on BRUMO and MATH-500, which still show meaningful separation.
Known limitations
AIME and HMMT are effectively solved by AI — they are displayed for reference but no longer factor into the weighted score. If math reasoning is critical for your use case, look at BRUMO scores specifically, and consider models with explicit reasoning capabilities (chain-of-thought). See the AIME & HMMT explainer.
How we weight
Mathematics carries a 5% weight in BenchLM.ai's overall scoring. Frontier models score 95-99% on AIME and HMMT — competition math is effectively solved by AI.
AIME and HMMT are still displayed for reference but no longer factor into the weighted score due to saturation. BRUMO and MATH-500 show more meaningful separation. If mathematical reasoning is critical, prioritize models with explicit reasoning capabilities. See the math leaderboard or read the AIME & HMMT explainer.
Leaderboards exclude benchmark rows that BenchLM generated from other scores or cloned from reference models. When a weighted benchmark is missing after that filter, the category falls back to the remaining trustworthy public rows instead of filling the gap with synthetic values.
The full scoring rules, freshness handling, and runtime/pricing caveats live on the BenchLM methodology page.
Scroll horizontally to read the full evidence ledger.
| Benchmark | Weight | Status | Description |
|---|---|---|---|
| AIME 2023 | — | Display only | High school mathematics competition |
| AIME 2024 | — | Display only | High school mathematics competition |
| AIME 2025 | — | Display only | High school mathematics competition |
| AIME25 (Arcee) | — | Display only | Display-only AIME25 reference from Arcee AI's Trinity-Large-Thinking launch chart. |
| HMMT Feb 2023 | — | Display only | Collegiate mathematics competition |
| HMMT Feb 2024 | — | Display only | Collegiate mathematics competition |
| HMMT Feb 2025 | — | Display only | Collegiate mathematics competition |
| BRUMO 2025 | — | Display only | University-level mathematics olympiad |
| MATH-500 | — | Display only | Curated 500-problem subset of the MATH dataset covering algebra, geometry, number theory, and more |
About Math Benchmarks
High school mathematics competition
Common questions
What is the best LLM for math?
The best LLMs for math are ranked by competition-level benchmarks like AIME and HMMT, with top models achieving strong scores on problems from real math competitions.
How are math benchmarks scored for LLMs?
Math benchmarks score models on correctness of final answers to problems ranging from algebra to advanced competition mathematics, testing both calculation and mathematical reasoning.
What benchmarks test math ability in AI models?
Major math benchmarks include AIME (competition-level problems), HMMT (Harvard-MIT tournament problems), and BRUMO, each testing progressively harder mathematical reasoning.
Math benchmark updates
Math model rankings change weekly. Stay current.
One email each week. Unsubscribe anytime.