# Best AI model for math — September 2026

> Best AI model for math (September 2026): on BenchLM's evidence, Kimi K2.6 has the highest math score estimate (71.0) among models that meet the page's constraints, followed by Claude Opus 4.8 and GLM-5.1. Evidence, constraints, and caveats included.

## Constraints on this page

A builder choosing by accuracy, hosted processing allowed, any price, ordinary input size. Every refine link below changes one answer.

## The shortlist for math

| # | Model | Creator | Math score estimate | Why |
|---|---|---|---|---|
| 1 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 71.0 | Highest math estimate among models that meet every stated constraint; $1.75 for the stated workload; 256K context; Open weights, so it can run on your own hardware |
| 2 | [Claude Opus 4.8](/models/claude-opus-4-8) | Anthropic | 65.2 | Estimate 65.2 on the same math evidence; $10.00 for the stated workload; 1M context |
| 3 | [GLM-5.1](/models/glm-5-1) | Z.AI | 63.8 | Estimate 63.8 on the same math evidence; $2.28 for the stated workload; 203K context; Open weights, so it can run on your own hardware |
| 4 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 62.1 | Estimate 62.1 on the same math evidence; $1.20 for the stated workload; 256K context; Open weights, so it can run on your own hardware |
| 5 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 62.0 | Estimate 62.0 on the same math evidence; 1M context |

Ordered by task score under the stated constraints.

## What this shortlist rests on

- 30% · [FrontierMath v2 (Tiers 1-3)](/benchmarks/frontiermathv2tiers13) · Current
- 25% · [AIME26](/benchmarks/aime2026) · Current
- 25% · [HMMT Feb 2026](/benchmarks/hmmtfeb2026) · Current
- 10% · [FrontierMath v2 (Tier 4)](/benchmarks/frontiermathv2tier4) · Current
- 10% · [USAMO 2026](/benchmarks/usamo2026) · Current

## What to verify before choosing

- Math tests are relevant evidence; tool execution and scientific judgment still need separate validation.
- Composite scores are estimates. A small score gap does not establish a reliably better model.

## Refine

- [Change any answer in the selector](/tools/llm-selector?audience=builders&family=learning&useCase=math&priority=accuracy&privacy=hosted&budget=any&context=normal)
- [Must run on my own hardware](/tools/llm-selector?audience=builders&family=learning&useCase=math&priority=accuracy&privacy=local&context=normal)
- [Cost matters most](/tools/llm-selector?audience=builders&family=learning&useCase=math&priority=cost&privacy=hosted&budget=any&context=normal)
- [Speed matters most](/tools/llm-selector?audience=builders&family=learning&useCase=math&priority=speed&privacy=hosted&budget=any&context=normal)
- [Long documents or a large codebase](/tools/llm-selector?audience=builders&family=learning&useCase=math&priority=accuracy&privacy=hosted&budget=any&context=large)
- [Choosing an everyday assistant, not an API](/tools/llm-selector?audience=everyday&family=learning&useCase=math&priority=accuracy&privacy=hosted&budget=any&context=normal)
- [Describe your own job instead](/tools/llm-selector)

## Questions

**Which AI model is best for math?**

On BenchLM's public evidence, Kimi K2.6 by Moonshot AI has the highest math score estimate among models that meet the page's default constraints (71.0). Ordered by task score under the stated constraints. Math tests are relevant evidence; tool execution and scientific judgment still need separate validation.

**What are the alternatives to Kimi K2.6 for math?**

Claude Opus 4.8 (65.2), GLM-5.1 (63.8), Kimi K2.5 (62.1), Qwen3.6 Plus (62.0) follow on the same evidence. A small gap does not establish a reliably better model; compare them on three representative examples of your own work.

**How does BenchLM pick the best ai model for math?**

The page runs the LLM Selector with fixed answers: a builder choosing by accuracy, hosted processing allowed, any price, ordinary input size. The selector uses the math evidence surface, filters by the stated constraints, and orders by that evidence. It never adds a hidden fit score or a bonus for open weights or reasoning style.

**Can I change the constraints?**

Yes. Every link under "Refine" opens the selector with one answer changed, and the address carries the answers so a result can be shared or reopened against the current dataset.

Data updated 2026-09-18. Method: bench-align-v5.5-2026-09-04.

Canonical page: https://benchlm.ai/best/for/math
