Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes

Instruction Following Benchmark (IFBench)

IFBench evaluates precise instruction-following generalization on 58 challenging, verifiable out-of-domain constraints. Unlike IFEval which tests familiar constraint types, IFBench specifically measures how well models follow novel instructions they haven't been optimized for, exposing overfitting to common instruction patterns.

Data verified 33 confirmed releases in the last 30 daysSee provider release alerts

MAI-Thinking-1 leads the IFBench leaderboard on BenchLM's September 2026 update with 85%, ahead of Qwen3.8 Max (82.8%) and Inkling-Small (82.2%), across 41 models.

Top models on IFBench — September 15, 2026

As of September 15, 2026, MAI-Thinking-1 leads the IFBench leaderboard with 85% , followed by Qwen3.8 Max (82.8%) and Inkling-Small (82.2%).

41 modelsInstruction Following70% of category scoreCurrentUpdated September 15, 2026

Leaderboard (41 models)

Score
1
MAI-Thinking-1Microsoft · Closed
85%
2
Qwen3.8 MaxAlibaba · Open weight
82.8%
3
Inkling-SmallThinking Machines Lab · Open weight
82.2%
4
Nemotron 3 UltraNVIDIA · Open weight
81.7%
5
Grok 4.3xAI · Closed
81.3%
6
Qwen3.8-Flash-NextAlibaba · Open weight
81.3%
7
dots3-note PreviewDots Studio · Open weight
80.4%
8
Solar Open 2Upstage · Open weight
80%
9
InklingThinking Machines Lab · Open weight
79.8%
10
Qwen3.8-27BAlibaba · Open weight
79.5%
11
Granite 4.2 8BIBM · Open weight
79.3%
12
Qwen3.7 MaxAlibaba · Closed
79.1%
13
Qwen3.7 PlusAlibaba · Closed
79.1%
14
Granite 4.2 30BIBM · Open weight
77.2%
15
Muse Glimmer 30BMeta · Open weight
77%
16
Mercury 2.5Inception · Closed
77%
17
Gemini 3.5 FlashGoogle · Closed
76.3%
18
A.X K2SK Telecom · Open weight
75.9%
19
Qwen3.6 PlusAlibaba · Closed
75.8%
20
Ling 3.0 FlashInclusionAI · Open weight
74.5%
21
Granite 4.2 3BIBM · Open weight
74.3%
22
Nemotron 3 Nano Omni 30B A3BNVIDIA · Open weight
74.2%
23
Ling 3.0 Flash FP8InclusionAI · Open weight
73.4%
24
72.9%
25
K-EXAONE 2.0LG AI Research · Open weight
72.6%
26
Mercury 2Inception · Closed
71.3%
27
Agents-A1-4BInternScience · Open weight
69.1%
28
MiniCPM5-2BOpenBMB · Open weight
66.3%
29
Hy3 PreviewTencent · Open weight
63.1%
30
LFM2.5-2.6BLiquidAI · Open weight
59.2%
31
Claude Opus 4.5Anthropic · Closed
58%
32
Ling 2.6 FlashInclusionAI · Open weight
57%
33
LFM2.5-8B-A1BLiquidAI · Open weight
56.5%
34
Solar Pro 3Upstage · Closed
55.8%
35
ZAYA1-8BZyphra · Open weight
52.6%
36
MiniCPM5-1BOpenBMB · Open weight
46.7%
37
LFM2.5-230MLiquidAI · Open weight
38.4%
38
Kanana-2 1.3B InstructKakao · Open weight
34.7%
39
Kanana-2 3B InstructKakao · Open weight
33.3%
40
LFM2.5-VL-3BLiquidAI · Open weight
25.8%
41
LLaDA2.2-miniInclusionAI · Open weight
24.9%

According to BenchLM.ai, MAI-Thinking-1 leads the IFBench benchmark with a score of 85%, followed by Qwen3.8 Max (82.8%) and Inkling-Small (82.2%). The top models are clustered within 2.8 points, suggesting this benchmark is nearing saturation for frontier models.

41 models have been evaluated on IFBench. The benchmark falls in the Instruction Following category. This category carries a 5% weight in BenchLM.ai's overall scoring system. Within that category, IFBench contributes 70% of the category score, so strong performance here directly affects a model's overall ranking.

About IFBench

Year

2025

Tasks

58

BenchLM freshness & provenance

Version

IFBench 2025

Refresh cadence

Quarterly

Staleness state

Current

Question availability

Public benchmark set

Current

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does IFBench measure?

IFBench evaluates precise instruction-following generalization on 58 challenging, verifiable out-of-domain constraints. Unlike IFEval which tests familiar constraint types, IFBench specifically measures how well models follow novel instructions they haven't been optimized for, exposing overfitting to common instruction patterns.

Which model scores highest on IFBench?

MAI-Thinking-1 by Microsoft currently leads with a score of 85% on IFBench.

How many models are evaluated on IFBench?

41 AI models have been evaluated on IFBench on BenchLM.

Last updated: September 15, 2026 · BenchLM version IFBench 2025

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.