FrontierMath v2 Tier 4 (FrontierMath v2 (Tier 4))
Epoch AI's corrected v2 Tier 4 expansion, a separate set of exceptionally difficult research-level mathematics problems evaluated with Python-enabled iterative reasoning.
Data verified 31 confirmed releases in the last 30 daysSee provider release alertsTop models on FrontierMath v2 (Tier 4) — September 10, 2026
As of September 10, 2026, GPT-6 Astra leads the FrontierMath v2 (Tier 4) leaderboard with 97.600% , followed by GPT-5.6 Sol (83.000%) and GPT-5.6 Terra (68.300%).
GPT-6 Astra
OpenAI
GPT-5.6 Sol
OpenAI
GPT-5.6 Terra
OpenAI
48 modelsMath10% of category scoreCurrentUpdated September 10, 2026
Leaderboard (48 models)
ScoreAccording to BenchLM.ai, GPT-6 Astra leads the FrontierMath v2 (Tier 4) benchmark with a score of 97.600%, followed by GPT-5.6 Sol (83.000%) and GPT-5.6 Terra (68.300%). There is significant spread across the leaderboard, making this benchmark effective at differentiating model capabilities.
48 models have been evaluated on FrontierMath v2 (Tier 4). The benchmark falls in the Math category. This category carries a 5% weight in BenchLM.ai's overall scoring system. Within that category, FrontierMath v2 (Tier 4) contributes 10% of the category score, so strong performance here directly affects a model's overall ranking.
About FrontierMath v2 (Tier 4)
Year
2026
Tasks
43 private extreme-difficulty mathematics problems
Format
Python-enabled iterative mathematical problem solving
Difficulty
Research-level mathematics requiring hours or days of expert work
The v2 private Tier 4 set contains 43 problems. Its smaller sample makes individual scores noisier than Tiers 1-3, so it is a separate ranking factor with lower weight rather than being averaged into the core score.
BenchLM freshness & provenance
Version
FrontierMath v2 (Tier 4) 2026
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does FrontierMath v2 (Tier 4) measure?
Epoch AI's corrected v2 Tier 4 expansion, a separate set of exceptionally difficult research-level mathematics problems evaluated with Python-enabled iterative reasoning.
Which model scores highest on FrontierMath v2 (Tier 4)?
GPT-6 Astra by OpenAI currently leads with a score of 97.600% on FrontierMath v2 (Tier 4).
How many models are evaluated on FrontierMath v2 (Tier 4)?
48 AI models have been evaluated on FrontierMath v2 (Tier 4) on BenchLM.
Compare Top Models on FrontierMath v2 (Tier 4)
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.