Skip to main content
Radar

Keep up with the models you depend on. Follow price changes, retirements, and API updates.Follow the models you depend on.

Follow model changes

FrontierMath v2 Tiers 1-3 (FrontierMath v2 (Tiers 1-3))

Epoch AI's corrected v2 core FrontierMath suite of private advanced mathematics problems. Models can reason iteratively and use Python; scores are pass rates on the private set.

Data verified 31 confirmed releases in the last 30 daysSee provider release alerts

Top models on FrontierMath v2 (Tiers 1-3) — September 10, 2026

As of September 10, 2026, GPT-5.6 Sol leads the FrontierMath v2 (Tiers 1-3) leaderboard with 89.000% , followed by GPT-5.6 Terra (84.900%) and GPT-5.6 Luna (78.600%).

53 modelsMath30% of category scoreCurrentUpdated September 10, 2026

Leaderboard (53 models)

Score
1
GPT-5.6 SolOpenAI · Closed
89.000%
2
GPT-5.6 TerraOpenAI · Closed
84.900%
3
GPT-5.6 LunaOpenAI · Closed
78.600%
4
GPT-5.5OpenAI · Closed
51.700%
5
GPT-5.5 ProOpenAI · Closed
51.000%
6
GPT-5.4 ProOpenAI · Closed
50.000%
7
GPT-5.4OpenAI · Closed
47.600%
8
Claude Opus 4.8Anthropic · Closed
47.241%
9
Claude Opus 4.7Anthropic · Closed
43.793%
10
Claude Opus 4.6Anthropic · Closed
40.700%
11
GPT-5.2OpenAI · Closed
40.700%
12
Muse SparkMeta · Closed
39.000%
13
Gemini 3.5 FlashGoogle · Closed
38.966%
14
Kimi K2.6Moonshot AI · Open weight
38.966%
15
Gemini 3 ProGoogle · Closed
37.600%
16
Gemini 3.1 ProGoogle · Closed
36.900%
17
Gemini 3 FlashGoogle · Closed
35.640%
18
GLM-5.1Z.AI · Open weight
33.448%
19
Claude Sonnet 4.6Anthropic · Closed
32.400%
20
GPT-5.1OpenAI · Closed
31.034%
21
GPT-5.4 miniOpenAI · Closed
28.280%
22
Kimi K2.5Moonshot AI · Open weight
27.900%
23
GPT-5 miniOpenAI · Closed
27.241%
24
Qwen3.6 PlusAlibaba · Closed
26.207%
25
GPT-5.4 nanoOpenAI · Closed
25.860%
26
o4-mini (high)OpenAI · Closed
24.828%
27
Qwen 3.6 Max (preview)Alibaba · Closed
23.103%
28
DeepSeek V3.2DeepSeek · Open weight
22.100%
29
Kimi K2Moonshot AI · Closed
21.404%
30
Qwen3.5 PlusAlibaba · Closed
21.034%
31
Claude Opus 4.5Anthropic · Closed
20.690%
32
Grok 4xAI · Closed
19.655%
33
o3OpenAI · Closed
18.685%
34
GLM-5Z.AI · Open weight
16.434%
35
Gemini 2.5 ProGoogle · Closed
14.138%
36
Claude Sonnet 4.5Anthropic · Closed
13.495%
37
o1OpenAI · Closed
9.310%
38
Qwen3 235B 2507 (Reasoning)Alibaba · Open weight
8.481%
39
GPT-5 nanoOpenAI · Closed
8.276%
40
Qwen3.5 FlashAlibaba · Closed
6.207%
41
Claude Haiku 4.5Anthropic · Closed
5.903%
42
GPT-4.1OpenAI · Closed
5.517%
43
Gemini 2.5 FlashGoogle · Closed
4.844%
44
GPT-4.1 miniOpenAI · Closed
4.483%
45
GLM-4.6Z.AI · Open weight
3.819%
46
Grok 3 [Beta]xAI · Closed
3.793%
47
GLM-4.7Z.AI · Open weight
2.439%
48
Claude 3.5 SonnetAnthropic · Closed
2.069%
49
DeepSeek V3DeepSeek · Open weight
1.724%
50
GPT-4.1 nanoOpenAI · Closed
1.034%
51
Llama 4 MaverickMeta · Open weight
0.690%
52
GPT-4oOpenAI · Closed
0.345%
53
Llama 4 ScoutMeta · Open weight
0.000%

According to BenchLM.ai, GPT-5.6 Sol leads the FrontierMath v2 (Tiers 1-3) benchmark with a score of 89.000%, followed by GPT-5.6 Terra (84.900%) and GPT-5.6 Luna (78.600%). There is significant spread across the leaderboard, making this benchmark effective at differentiating model capabilities.

53 models have been evaluated on FrontierMath v2 (Tiers 1-3). The benchmark falls in the Math category. This category carries a 5% weight in BenchLM.ai's overall scoring system. Within that category, FrontierMath v2 (Tiers 1-3) contributes 30% of the category score, so strong performance here directly affects a model's overall ranking.

About FrontierMath v2 (Tiers 1-3)

Year

2026

Tasks

295 private advanced mathematics problems

Format

Python-enabled iterative mathematical problem solving

Difficulty

From olympiad-plus to early research mathematics

After Epoch AI's June 2026 correction, the FrontierMath v2 private core contains 295 Tiers 1-3 problems. BenchLM selects the highest published reasoning-effort result for each model and keeps this core score distinct from Tier 4.

BenchLM freshness & provenance

Version

FrontierMath v2 (Tiers 1-3) 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does FrontierMath v2 (Tiers 1-3) measure?

Epoch AI's corrected v2 core FrontierMath suite of private advanced mathematics problems. Models can reason iteratively and use Python; scores are pass rates on the private set.

Which model scores highest on FrontierMath v2 (Tiers 1-3)?

GPT-5.6 Sol by OpenAI currently leads with a score of 89.000% on FrontierMath v2 (Tiers 1-3).

How many models are evaluated on FrontierMath v2 (Tiers 1-3)?

53 AI models have been evaluated on FrontierMath v2 (Tiers 1-3) on BenchLM.

Last updated: September 10, 2026 · BenchLM version FrontierMath v2 (Tiers 1-3) 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.