Skip to main content

Benchmark profile

American Invitational Mathematics Examination 2025 (AIME 2025)

The most recent AIME examination, featuring 15 challenging mathematics problems testing olympiad-level mathematical reasoning with integer answers from 000-999.

Data verified

Benchmark score on AIME 2025 — July 20, 2026

BenchLM mirrors the published score view for AIME 2025. MAI-Thinking-1 leads the public snapshot at 97% , followed by Kimi K2.5 (96.1%) and Kimi K2.5 (Reasoning) (96.1%). BenchLM does not use these results to rank models overall.

11 modelsMathCurrentDisplay onlyUpdated July 20, 2026

Benchmark score table (11 models)

Score
1
MAI-Thinking-1Microsoft · Closed
97%
2
Kimi K2.5Moonshot AI · Open weight
96.1%
3
Kimi K2.5 (Reasoning)Moonshot AI · Closed
96.1%
4
GLM-4.7Z.AI · Open weight
95.7%
5
MiMo-V2-FlashXiaomi · Open weight
94.1%
6
Claude Sonnet 4.5Anthropic · Closed
87%
7
Exaone 4.0 32BLG AI Research · Open weight
85.3%
8
GPT-5 nanoOpenAI · Closed
85.2%
9
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
82.1%
10
LFM2.5-8B-A1BLiquidAI · Open weight
42.5%
11
MiniCPM5-1BOpenBMB · Open weight
40.4%

The published AIME 2025 snapshot places MAI-Thinking-1 first at 97%. The third row is 0.9 points behind. The broader top-10 range is 54.5 points, so the table still separates the published systems.

11 models have been evaluated on AIME 2025. The benchmark falls in the Math category. This category carries a 5% weight in BenchLM.ai's overall scoring system. AIME 2025 is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About AIME 2025

Year

2025

Tasks

15 problems

Format

Integer answers 000-999

Difficulty

High school olympiad level

AIME 2025 represents the current standard for intermediate-level mathematical olympiad problems. Success requires sophisticated mathematical reasoning and problem-solving techniques.

BenchLM freshness & provenance

Version

AIME 2025

Refresh cadence

Annual

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does AIME 2025 measure?

The most recent AIME examination, featuring 15 challenging mathematics problems testing olympiad-level mathematical reasoning with integer answers from 000-999.

Which model scores highest on AIME 2025?

MAI-Thinking-1 by Microsoft currently leads with a score of 97% on AIME 2025.

How many models are evaluated on AIME 2025?

11 AI models have been evaluated on AIME 2025 on BenchLM.

Last updated: July 20, 2026 · BenchLM version AIME 2025

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.