# AIME 2026 (AIME26)

> A 2026 American Invitational Mathematics Examination snapshot used in frontier-model comparison tables for mathematical reasoning.

Canonical page: https://benchlm.ai/benchmarks/aime2026

- Category: [Mathematics](/math)
- Last updated: October 7, 2026

## About AIME26

- Year: 2026
- Tasks: Competition math problems
- Format: Short-answer mathematics
- Difficulty: Olympiad-style mathematics
- Paper: [Qwen3.6 launch benchmarks](https://qwen.ai/blog?id=qwen3.6)

AIME-style benchmarks remain one of the fastest ways to separate top reasoning models on olympiad-style math. AIME 2026 is a newer contest-year snapshot than the legacy AIME rows already tracked on BenchLM.

The Mathematics leaderboard ranks models by a weighted category score, and AIME26 contributes 25% of it. BenchAlign v5.8 also uses it in the overall ranking.

## Leaderboard (29 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [GLM-5.2](/models/glm-5-2) | Z.AI | 99.2% |
| 2 | [Beam](/models/reflection-beam) | Reflection AI | 97.8% |
| 3 | [Inkling](/models/inkling) | Thinking Machines Lab | 97.1% |
| 4 | [A.X K2](/models/a-x-k2) | SK Telecom | 97.1% |
| 5 | [Kimi K2.6](/models/kimi-2-6) | Moonshot AI | 96.4% |
| 6 | [Ternary Bonsai 2 27B](/models/ternary-bonsai-2-27b) | Prism ML | 95.8% |
| 7 | [GLM-5](/models/glm-5) | Z.AI | 95.8% |
| 8 | [Kimi K2.5](/models/kimi-k2-5) | Moonshot AI | 95.8% |
| 9 | [Solar Open 2](/models/solar-open-2) | Upstage | 95.7% |
| 10 | [Inkling-Small](/models/inkling-small) | Thinking Machines Lab | 95.5% |
| 11 | [GLM-5.1](/models/glm-5-1) | Z.AI | 95.3% |
| 12 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 95.3% |
| 13 | [Solar Pro 4](/models/solar-pro-4) | Upstage | 95.3% |
| 14 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 95.1% |
| 15 | [Muse Glimmer 30B](/models/muse-glimmer-30b) | Meta | 94.7% |
| 16 | [MAI-Thinking-1](/models/mai-thinking-1) | Microsoft | 94.5% |
| 17 | [Qwen3.6-27B](/models/qwen3-6-27b) | Alibaba | 94.1% |
| 18 | [Qwen3.5 397B](/models/qwen3-5-397b) | Alibaba | 93.3% |
| 19 | [Ling 3.0 Flash](/models/ling-3-0-flash) | InclusionAI | 93.2% |
| 20 | [Qwen3.6-35B-A3B](/models/qwen3-6-35b-a3b) | Alibaba | 92.7% |
| 21 | [K-EXAONE 2.0](/models/k-exaone-2-0) | LG AI Research | 92.3% |
| 22 | [ZAYA1-8B](/models/zaya1-8b) | Zyphra | 89.1% |
| 23 | [MiniCPM5-2B](/models/minicpm5-2b) | OpenBMB | 86.5% |
| 24 | [Gemma 4 12B](/models/gemma-4-12b) | Google | 77.5% |
| 25 | [ZAYA1-74B-Preview](/models/zaya1-74b-preview) | Zyphra | 76.4% |
| 26 | [LongCat-Flash-Lite-Sparse](/models/longcat-flash-lite-sparse) | Meituan | 65.7% |
| 27 | [LFM2.5-8B-A1B](/models/lfm2-5-8b-a1b) | LiquidAI | 50.0% |
| 28 | [MiniCPM5-1B](/models/minicpm5-1b) | OpenBMB | 40.4% |
| 29 | [LLaDA2.2-mini](/models/llada2-2-mini) | InclusionAI | 35.0% |

## FAQ

### What does AIME26 measure?

A 2026 American Invitational Mathematics Examination snapshot used in frontier-model comparison tables for mathematical reasoning.

### Which model scores highest on AIME26?

GLM-5.2 by Z.AI currently leads with a score of 99.2% on AIME26.

### How many models are evaluated on AIME26?

29 AI models have published results on AIME26 in the BenchLM catalog.

### Does AIME26 affect BenchLM's overall score?

Yes. The Mathematics leaderboard ranks models by a weighted category score, and AIME26 contributes 25% of it. BenchAlign v5.8 also uses it in the overall ranking.

## Compare Top Models on AIME26

- [GLM-5.2 vs Beam](/compare/glm-5-2-vs-reflection-beam)
- [Beam vs Inkling](/compare/inkling-vs-reflection-beam)
- [Inkling vs A.X K2](/compare/a-x-k2-vs-inkling)
- [A.X K2 vs Kimi K2.6](/compare/a-x-k2-vs-kimi-2-6)
