Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes

Harvard-MIT Mathematics Tournament February 2026 (HMMT Feb 2026)

A February 2026 HMMT slice used in newer frontier-model math comparisons.

Data verified 33 confirmed releases in the last 30 daysSee provider release alerts

Top models on HMMT Feb 2026 — September 15, 2026

As of September 15, 2026, Qwen3.7 Max leads the HMMT Feb 2026 leaderboard with 97.1% , followed by DeepSeek V4 Pro 0813 (95.2%) and DeepSeek V4 Flash 0731 (94.8%).

27 modelsMath25% of category scoreCurrentUpdated September 15, 2026

Leaderboard (27 models)

Score
1
Qwen3.7 MaxAlibaba · Closed
97.1%
2
DeepSeek V4 Pro 0813DeepSeek · Closed
95.2%
3
DeepSeek V4 Flash 0731DeepSeek · Closed
94.8%
4
DeepSeek V4 Pro (High)DeepSeek · Open weight
94.0%
5
Solar Open 2Upstage · Open weight
93.9%
6
Qwen3.7 PlusAlibaba · Closed
92.9%
7
Kimi K2.6Moonshot AI · Open weight
92.7%
8
GLM-5.2Z.AI · Open weight
92.5%
9
DeepSeek V4 Flash (High)DeepSeek · Closed
91.9%
10
Inkling-SmallThinking Machines Lab · Open weight
90.2%
11
Qwen3.5 397BAlibaba · Open weight
87.9%
12
Qwen3.6 PlusAlibaba · Closed
87.8%
13
Kimi K2.5Moonshot AI · Open weight
87.1%
14
Ling 3.0 FlashInclusionAI · Open weight
87.0%
15
GLM-5Z.AI · Open weight
86.4%
16
Claude Opus 4.5Anthropic · Closed
85.3%
17
MAI-Thinking-1Microsoft · Closed
84.9%
18
Qwen3.6-27BAlibaba · Open weight
84.3%
19
Qwen3.6-35B-A3BAlibaba · Open weight
83.6%
20
GLM-5.1Z.AI · Open weight
82.6%
21
K-EXAONE 2.0LG AI Research · Open weight
78.4%
22
ZAYA1-8BZyphra · Open weight
71.6%
23
MiniCPM5-2BOpenBMB · Open weight
63.8%
24
DeepSeek V4 FlashDeepSeek · Closed
40.8%
25
LongCat-Flash-Lite-SparseMeituan · Open weight
40.5%
26
DeepSeek V4 ProDeepSeek · Open weight
31.7%
27
MiniCPM5-1BOpenBMB · Open weight
25.8%

According to BenchLM.ai, Qwen3.7 Max leads the HMMT Feb 2026 benchmark with a score of 97.1%, followed by DeepSeek V4 Pro 0813 (95.2%) and DeepSeek V4 Flash 0731 (94.8%). The top models are clustered within 2.3 points, suggesting this benchmark is nearing saturation for frontier models.

27 models have been evaluated on HMMT Feb 2026. The benchmark falls in the Math category. This category carries a 5% weight in BenchLM.ai's overall scoring system. Within that category, HMMT Feb 2026 contributes 25% of the category score, so strong performance here directly affects a model's overall ranking.

About HMMT Feb 2026

Year

2026

Tasks

Competition math problems

Format

Contest mathematics

Difficulty

Olympiad-style mathematics

HMMT February 2026 matters because small score deltas at the frontier often depend on which contest set is used. BenchLM keeps this newer slice distinct from older HMMT summary rows.

BenchLM freshness & provenance

Version

HMMT Feb 2026 2026

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does HMMT Feb 2026 measure?

A February 2026 HMMT slice used in newer frontier-model math comparisons.

Which model scores highest on HMMT Feb 2026?

Qwen3.7 Max by Alibaba currently leads with a score of 97.1% on HMMT Feb 2026.

How many models are evaluated on HMMT Feb 2026?

27 AI models have been evaluated on HMMT Feb 2026 on BenchLM.

Last updated: September 15, 2026 · BenchLM version HMMT Feb 2026 2026

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.