Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes
Learn, research, or solve a problem · Math and scientific calculations

Best AI model for math — September 2026

On BenchLM's public evidence, Kimi K2.6 has the highest math score estimate among the models that meet this page's constraints (71.0). Math tests are relevant evidence; tool execution and scientific judgment still need separate validation.

Shortlist data updated

Share this shortlist

Share on XLinkedIn

This address is permanent. It always shows the current shortlist for this job, with the date the data was last updated. The embed shows the same shortlist on your site and links back here.

The shortlist for math

Constraints on this page: a builder choosing by accuracy, hosted processing allowed, any price, ordinary input size. Change any of them under Refine. A small score gap does not establish a reliably better model.

Compare Kimi K2.6 vs Claude Opus 4.8Full math leaderboard

  1. 01Kimi K2.6Best fit

    Moonshot AI · Open Weight

    71.0

    Math score estimate

    • Highest math estimate among models that meet every stated constraint
    • $1.75 for the stated workload
    • 256K context
    • Open weights, so it can run on your own hardware

    Model page and evidence

  2. 02Claude Opus 4.8

    Anthropic · Proprietary

    65.2

    Math score estimate

    • Estimate 65.2 on the same math evidence
    • $10.00 for the stated workload
    • 1M context

    Model page and evidenceCompare with Kimi K2.6

  3. 03GLM-5.1

    Z.AI · Open Weight

    63.8

    Math score estimate

    • Estimate 63.8 on the same math evidence
    • $2.28 for the stated workload
    • 203K context
    • Open weights, so it can run on your own hardware

    Model page and evidenceCompare with Kimi K2.6

  4. 04Kimi K2.5

    Moonshot AI · Open Weight

    62.1

    Math score estimate

    • Estimate 62.1 on the same math evidence
    • $1.20 for the stated workload
    • 256K context
    • Open weights, so it can run on your own hardware

    Model page and evidenceCompare with Kimi K2.6

  5. 05Qwen3.6 Plus

    Alibaba · Proprietary

    62.0

    Math score estimate

    • Estimate 62.0 on the same math evidence
    • 1M context

    Model page and evidenceCompare with Kimi K2.6

Refine for your situation

Each link opens the selector with one answer changed. The address carries the answers, so your version is as shareable as this page.

A decision model maps your sentence to the selector’s questions. The shortlist comes from public evidence only, and nothing is stored.

What this shortlist rests on

The math surface. The category score is a weighted average of these public benchmarks.

What to verify before choosing

  • Math tests are relevant evidence; tool execution and scientific judgment still need separate validation.
  • Composite scores are estimates. A small score gap does not establish a reliably better model.

Try three representative examples of your own work. Compare errors, time, cost, and the tools available in your actual setup.

Questions

Which AI model is best for math?

On BenchLM's public evidence, Kimi K2.6 by Moonshot AI has the highest math score estimate among models that meet the page's default constraints (71.0). Ordered by task score under the stated constraints. Math tests are relevant evidence; tool execution and scientific judgment still need separate validation.

What are the alternatives to Kimi K2.6 for math?

Claude Opus 4.8 (65.2), GLM-5.1 (63.8), Kimi K2.5 (62.1), Qwen3.6 Plus (62.0) follow on the same evidence. A small gap does not establish a reliably better model; compare them on three representative examples of your own work.

How does BenchLM pick the best ai model for math?

The page runs the LLM Selector with fixed answers: a builder choosing by accuracy, hosted processing allowed, any price, ordinary input size. The selector uses the math evidence surface, filters by the stated constraints, and orders by that evidence. It never adds a hidden fit score or a bonus for open weights or reasoning style.

Can I change the constraints?

Yes. Every link under "Refine" opens the selector with one answer changed, and the address carries the answers so a result can be shared or reopened against the current dataset.

Method: bench-align-v5.5-2026-09-04. Read the methodology and benchmark confidence pages for how scores and verification statuses are produced.

Watch the math shortlist

One weekly email when rank, price, or benchmark evidence changes make this shortlist worth revisiting.

Read a sample issue

Join 2,000+ readers.