AIME 2026 (AIME26)
A 2026 American Invitational Mathematics Examination snapshot used in frontier-model comparison tables for mathematical reasoning.
Top models on AIME26 — October 7, 2026
As of October 7, 2026, GLM-5.2 leads the AIME26 leaderboard with 99.2% , followed by Beam (97.8%) and Inkling (97.1%).
GLM-5.2
Z.AI
Beam
Reflection AI
Inkling
Thinking Machines Lab
29 modelsMathematics25% of category scoreCurrentUpdated October 7, 2026
| Rank | Model / configuration | Score | Parameters (B) | Open / closed |
|---|---|---|---|---|
| 1 | GLM-5.2Z.AI | 99.2% | Not reported | Open |
| 2 | BeamReflection AI | 97.8% | Not reported | Pending |
| 3 | InklingThinking Machines Lab | 97.1% | Not reported | Open |
| 4 | A.X K2SK Telecom | 97.1% | Not reported | Open |
| 5 | Kimi K2.6Moonshot AI | 96.4% | Not reported | Open |
| 6 | Ternary Bonsai 2 27BPrism ML | 95.8% | Not reported | Open |
| 7 | GLM-5Z.AI | 95.8% | Not reported | Open |
| 8 | Kimi K2.5Moonshot AI | 95.8% | Not reported | Open |
| 9 | Solar Open 2Upstage | 95.7% | Not reported | Open |
| 10 | Inkling-SmallThinking Machines Lab | 95.5% | Not reported | Open |
| 11 | GLM-5.1Z.AI | 95.3% | Not reported | Open |
| 12 | Qwen3.6 PlusAlibaba | 95.3% | Not reported | Closed |
| 13 | Solar Pro 4Upstage | 95.3% | Not reported | Closed |
| 14 | Claude Opus 4.5Anthropic | 95.1% | Not reported | Closed |
| 15 | Muse Glimmer 30BMeta | 94.7% | Not reported | Open |
| 16 | MAI-Thinking-1Microsoft | 94.5% | Not reported | Closed |
| 17 | Qwen3.6-27BAlibaba | 94.1% | Not reported | Open |
| 18 | Qwen3.5 397BAlibaba | 93.3% | Not reported | Open |
| 19 | Ling 3.0 FlashInclusionAI | 93.2% | Not reported | Open |
| 20 | Qwen3.6-35B-A3BAlibaba | 92.7% | Not reported | Open |
| 21 | K-EXAONE 2.0LG AI Research | 92.3% | Not reported | Open |
| 22 | ZAYA1-8BZyphra | 89.1% | Not reported | Open |
| 23 | MiniCPM5-2BOpenBMB | 86.5% | Not reported | Open |
| 24 | Gemma 4 12BGoogle | 77.5% | Not reported | Open |
| 25 | ZAYA1-74B-PreviewZyphra | 76.4% | Not reported | Open |
| 26 | LongCat-Flash-Lite-SparseMeituan | 65.7% | Not reported | Open |
| 27 | LFM2.5-8B-A1BLiquidAI | 50.0% | Not reported | Open |
| 28 | MiniCPM5-1BOpenBMB | 40.4% | Not reported | Open |
| 29 | LLaDA2.2-miniInclusionAI | 35.0% | Not reported | Open |
According to BenchLM.ai, GLM-5.2 leads the AIME26 benchmark with a score of 99.2%, followed by Beam (97.8%) and Inkling (97.1%). The top models are clustered within 2.1 points, suggesting this benchmark is nearing saturation for frontier models.
29 models have been evaluated on AIME26. The benchmark falls in the Mathematics category. The Mathematics leaderboard ranks models by a weighted category score, and AIME26 contributes 25% of it. BenchAlign v5.8 also uses it in the overall ranking.
About AIME26
Year
2026
Tasks
Competition math problems
Format
Short-answer mathematics
Difficulty
Olympiad-style mathematics
AIME-style benchmarks remain one of the fastest ways to separate top reasoning models on olympiad-style math. AIME 2026 is a newer contest-year snapshot than the legacy AIME rows already tracked on BenchLM.
Freshness and provenance
Version
AIME26 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
Questions
What does AIME26 measure?
A 2026 American Invitational Mathematics Examination snapshot used in frontier-model comparison tables for mathematical reasoning.
Which model scores highest on AIME26?
GLM-5.2 by Z.AI currently leads with a score of 99.2% on AIME26.
How many models are evaluated on AIME26?
29 AI models have published results on AIME26 in the BenchLM catalog.
Compare top models on AIME26
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 5,500+ readers.
One email each week. Unsubscribe anytime.