Skip to main content
BenchLM

Vals AIME (AIME)

We show this table for reference; we do not rank on it.

Data verified 36 confirmed releases in the last 30 daysFollow model changes

Challenging national math exam given to top high-school students

AIME score on AIME — April 16, 2026

We mirror the published aime score view for AIME. Gemini 3.1 Pro Preview leads the public snapshot at 98.13%, followed by GPT-5.2 (96.88%) and Muse Spark (96.88%). We do not use these results to rank models overall.

96 modelsMathematicsCurrentDisplay onlyUpdated April 16, 2026

AIME score table (96 models)

Score
1
Gemini 3.1 Pro PreviewGooglehigh reasoning
98.13%
2
GPT-5.2OpenAI · Closedxhigh reasoning
96.88%
3
Muse SparkMeta · Closed
96.88%
4
Gemini 3 Pro PreviewGooglehigh reasoning
96.68%
5
GPT-5.4OpenAI · Closedxhigh reasoning
96.67%
7
Claude Opus 4.7Anthropic · Closed
96.25%
8
Gemini 3 Flash PreviewGooglehigh reasoning
95.63%
9
Claude Opus 4.6 (Adaptive)Anthropic · Closed
95.63%
10
GPT-5.4 miniOpenAI · Closedxhigh reasoning
95.63%
11
95.63%
12
Claude Opus 4.5 ThinkingAnthropic · Closed
95.42%
13
Qwen3.6 PlusAlibaba · Closed
94.58%
14
GPT-5OpenAIhigh reasoning
93.37%
15
GPT-5.1OpenAI · Closedhigh reasoning
93.33%
16
GLM-4.7Z.AI · Open weight
93.33%
17
GLM-4.6Z.AI · Open weight
92.71%
18
GPT-OSS 120BOpenAI · Open weight
92.60%
19
Qwen3.5 FlashAlibaba · Closed
92.50%
20
Claude Sonnet 4.6Anthropic · Closed
92.29%
21
91.88%
22
GLM-5.1Z.AI · Open weight
91.88%
23
91.67%
24
GPT-5 miniOpenAI · Closedhigh reasoning
91.46%
25
91.25%
26
MiniMax M2.7MiniMax · Open weight
91.04%
27
90.56%
28
GPT-5.4 nanoOpenAI · Closedhigh reasoning
88.75%
29
MiniMax M2.5MiniMax · Closed
88.75%
30
Claude Sonnet 4.5 ThinkingAnthropic · Closed
88.19%
31
GLM-4.5Z.AI · Closed
86.67%
32
o3-miniOpenAI · Closedhigh reasoning
86.46%
33
86.04%
34
GPT-OSS 20BOpenAI · Open weight
86.04%
36
Kimi K2 ThinkingMoonshot AI
85.42%
37
o3OpenAI · Closedhigh reasoning
85.28%
38
85.00%
39
DeepSeek V3p2 ThinkingFireworks AIhigh reasoning
84.58%
40
Qwen3 235b A22bFireworks AI
83.96%
41
O4 MiniOpenAIhigh reasoning
83.67%
42
83.54%
43
Gemini 3.1 Flash Lite PreviewGooglehigh reasoning
83.33%
44
Claude Haiku 4.5 ThinkingAnthropic · Closed
82.71%
45
GPT-5 nanoOpenAI · Closedhigh reasoning
81.18%
46
Qwen3 MaxAlibaba · Closed
81.04%
47
80.68%
49
77.92%
50
Claude Opus 4.5Anthropic · Closed
76.88%
52
DeepSeek-R1DeepSeek · Open weight
73.96%
53
o1OpenAI · Closedhigh reasoning
71.46%
54
70.63%
55
DeepSeek V3p2Fireworks AInone reasoning
64.79%
56
62.71%
57
60.69%
58
58.75%
60
DeepSeek V3 0324Fireworks AI
52.20%
63
GPT-4.1 miniOpenAI · Closedhigh reasoning
49.38%
65
44.24%
66
42.92%
67
42.29%
69
Claude Opus 4Anthropic
41.25%
70
GPT-4.1OpenAI · Closedhigh reasoning
39.58%
71
38.54%
73
29.79%
74
DeepSeek V3DeepSeek · Open weight
27.50%
76
GPT-4.1 nanoOpenAI · Closedhigh reasoning
26.46%
78
25.21%
79
22.29%
81
18.75%
82
17.29%
84
15.21%
85
GPT-4oOpenAI · Closedhigh reasoning
13.96%
86
13.33%
87
GPT-4oOpenAI · Closedhigh reasoning
11.88%
88
GPT-4o miniOpenAI · Closedhigh reasoning
11.46%
89
Claude 3.5 SonnetAnthropic · Closed
10.00%
91
9.17%
92
5.63%
93
3.54%
94
3.33%
95
0.42%
96
0.42%

How AIME is shown here

BenchLM mirrors the public Vals AI AIME leaderboard captured from https://www.vals.ai/benchmarks/aime and updated by Vals on April 16, 2026. The snapshot preserves overall scores, uncertainty, latency, cost-per-test metadata, and task-level scores where Vals publishes them.

AIME is display only on BenchLM. Vals proprietary or Vals-hosted aggregate views are useful context, but BenchLM does not use them as weighted ranking inputs or as a replacement for benchmark-native source records.

Snapshot

96 Vals rows3 task viewspublic datasetTasks: Overall, AIME 2024, AIME 2025Display only

The published AIME snapshot places Gemini 3.1 Pro Preview first at 98.13%. The third row is 1.25 points behind. The broader top-10 range is 2.50 points, so many of the published results sit in a relatively narrow band.

96 models have been evaluated on AIME. The benchmark falls in the Mathematics category. We keep external benchmark mirrors separate from the weighted global scoring system, so these results remain source-specific evidence. AIME is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About AIME

Year

2026

Tasks

AIME math problems

Format

Accuracy score

Difficulty

Competition math

BenchLM mirrors the public Vals AI AIME leaderboard as display-only external evidence. The captured snapshot preserves overall scores, task-level scores where Vals publishes them, uncertainty, latency, and cost-per-test metadata. It is excluded from BenchLM weighted rankings.

Freshness and provenance

Version

AIME 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

CurrentDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

Questions

What does AIME measure?

Challenging national math exam given to top high-school students

Which model leads the published AIME snapshot?

Gemini 3.1 Pro Preview currently leads the published AIME snapshot with 98.13% aime score. BenchLM shows this benchmark for display only and does not use it in overall rankings.

How many models are evaluated on AIME?

The April 16, 2026 snapshot contains 96 AI models.

Last updated: April 16, 2026 · mirrored from the public benchmark leaderboard

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.