Instruction Following Benchmark (IFBench)
IFBench evaluates precise instruction-following generalization on 58 challenging, verifiable out-of-domain constraints. Unlike IFEval which tests familiar constraint types, IFBench specifically measures how well models follow novel instructions they haven't been optimized for, exposing overfitting to common instruction patterns.
Data verified 33 confirmed releases in the last 30 daysSee provider release alertsMAI-Thinking-1 leads the IFBench leaderboard on BenchLM's September 2026 update with 85%, ahead of Qwen3.8 Max (82.8%) and Inkling-Small (82.2%), across 41 models.
Top models on IFBench — September 15, 2026
As of September 15, 2026, MAI-Thinking-1 leads the IFBench leaderboard with 85% , followed by Qwen3.8 Max (82.8%) and Inkling-Small (82.2%).
MAI-Thinking-1
Microsoft
Qwen3.8 Max
Alibaba
Inkling-Small
Thinking Machines Lab
41 modelsInstruction Following70% of category scoreCurrentUpdated September 15, 2026
Leaderboard (41 models)
ScoreAccording to BenchLM.ai, MAI-Thinking-1 leads the IFBench benchmark with a score of 85%, followed by Qwen3.8 Max (82.8%) and Inkling-Small (82.2%). The top models are clustered within 2.8 points, suggesting this benchmark is nearing saturation for frontier models.
41 models have been evaluated on IFBench. The benchmark falls in the Instruction Following category. This category carries a 5% weight in BenchLM.ai's overall scoring system. Within that category, IFBench contributes 70% of the category score, so strong performance here directly affects a model's overall ranking.
About IFBench
Year
2025
Tasks
58
BenchLM freshness & provenance
Version
IFBench 2025
Refresh cadence
Quarterly
Staleness state
Current
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does IFBench measure?
IFBench evaluates precise instruction-following generalization on 58 challenging, verifiable out-of-domain constraints. Unlike IFEval which tests familiar constraint types, IFBench specifically measures how well models follow novel instructions they haven't been optimized for, exposing overfitting to common instruction patterns.
Which model scores highest on IFBench?
MAI-Thinking-1 by Microsoft currently leads with a score of 85% on IFBench.
How many models are evaluated on IFBench?
41 AI models have been evaluated on IFBench on BenchLM.
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.