# Instruction Following Benchmark (IFBench)

> IFBench evaluates precise instruction-following generalization on 58 challenging, verifiable out-of-domain constraints. Unlike IFEval which tests familiar constraint types, IFBench specifically measures how well models follow novel instructions they haven't been optimized for, exposing overfitting to common instruction patterns.

Canonical page: https://benchlm.ai/benchmarks/ifbench

- Category: [Instruction Following](/instruction-following)
- Last updated: September 15, 2026

## About IFBench

- Year: 2025
- Tasks: 58

IFBench is currently weighted in BenchLM's scoring formula. The Instruction Following category carries 5% of the overall score, and IFBench contributes 70% of that category score.

## Leaderboard (40 models)

| Rank | Model | Creator | Score |
|------|-------|---------|-------|
| 1 | [MAI-Thinking-1](/models/mai-thinking-1) | Microsoft | 85% |
| 2 | [Qwen3.8 Max](/models/qwen3-8-max) | Alibaba | 82.8% |
| 3 | [Inkling-Small](/models/inkling-small) | Thinking Machines Lab | 82.2% |
| 4 | [Nemotron 3 Ultra](/models/nemotron-3-ultra) | NVIDIA | 81.7% |
| 5 | [Grok 4.3](/models/grok-4-3) | xAI | 81.3% |
| 6 | [Qwen3.8-Flash-Next](/models/qwen3-8-flash-next) | Alibaba | 81.3% |
| 7 | [dots3-note Preview](/models/dots3-note-preview) | Dots Studio | 80.4% |
| 8 | [Solar Open 2](/models/solar-open-2) | Upstage | 80% |
| 9 | [Inkling](/models/inkling) | Thinking Machines Lab | 79.8% |
| 10 | [Qwen3.8-27B](/models/qwen3-8-27b) | Alibaba | 79.5% |
| 11 | [Granite 4.2 8B](/models/granite-4-2-8b) | IBM | 79.3% |
| 12 | [Qwen3.7 Max](/models/qwen3-7-max) | Alibaba | 79.1% |
| 13 | [Qwen3.7 Plus](/models/qwen3-7-plus) | Alibaba | 79.1% |
| 14 | [Granite 4.2 30B](/models/granite-4-2-30b) | IBM | 77.2% |
| 15 | [Muse Glimmer 30B](/models/muse-glimmer-30b) | Meta | 77% |
| 16 | [Mercury 2.5](/models/mercury-2-5) | Inception | 77% |
| 17 | [Gemini 3.5 Flash](/models/gemini-3-5-flash) | Google | 76.3% |
| 18 | [A.X K2](/models/a-x-k2) | SK Telecom | 75.9% |
| 19 | [Qwen3.6 Plus](/models/qwen3-6-plus) | Alibaba | 75.8% |
| 20 | [Ling 3.0 Flash](/models/ling-3-0-flash) | InclusionAI | 74.5% |
| 21 | [Granite 4.2 3B](/models/granite-4-2-3b) | IBM | 74.3% |
| 22 | [Nemotron 3 Nano Omni 30B A3B](/models/nemotron-3-nano-omni-30b-a3b) | NVIDIA | 74.2% |
| 23 | [Ling 3.0 Flash FP8](/models/ling-3-0-flash-fp8) | InclusionAI | 73.4% |
| 24 | [Nemotron 3.5 Lightning 30B A3B NVFP4](/models/nemotron-3-5-lightning-30b-a3b-nvfp4) | NVIDIA | 72.9% |
| 25 | [K-EXAONE 2.0](/models/k-exaone-2-0) | LG AI Research | 72.6% |
| 26 | [Agents-A1-4B](/models/agents-a1-4b) | InternScience | 69.1% |
| 27 | [MiniCPM5-2B](/models/minicpm5-2b) | OpenBMB | 66.3% |
| 28 | [Hy3 Preview](/models/hy3-preview) | Tencent | 63.1% |
| 29 | [LFM2.5-2.6B](/models/lfm2-5-2-6b) | LiquidAI | 59.2% |
| 30 | [Claude Opus 4.5](/models/claude-opus-4-5) | Anthropic | 58% |
| 31 | [Ling 2.6 Flash](/models/ling-2-6-flash) | InclusionAI | 57% |
| 32 | [LFM2.5-8B-A1B](/models/lfm2-5-8b-a1b) | LiquidAI | 56.5% |
| 33 | [Solar Pro 3](/models/solar-pro-3) | Upstage | 55.8% |
| 34 | [ZAYA1-8B](/models/zaya1-8b) | Zyphra | 52.6% |
| 35 | [MiniCPM5-1B](/models/minicpm5-1b) | OpenBMB | 46.7% |
| 36 | [LFM2.5-230M](/models/lfm2-5-230m) | LiquidAI | 38.4% |
| 37 | [Kanana-2 1.3B Instruct](/models/kanana-2-1-3b-instruct) | Kakao | 34.7% |
| 38 | [Kanana-2 3B Instruct](/models/kanana-2-3b-instruct) | Kakao | 33.3% |
| 39 | [LFM2.5-VL-3B](/models/lfm2-5-vl-3b) | LiquidAI | 25.8% |
| 40 | [LLaDA2.2-mini](/models/llada2-2-mini) | InclusionAI | 24.9% |

## FAQ

### What does IFBench measure?

IFBench evaluates precise instruction-following generalization on 58 challenging, verifiable out-of-domain constraints. Unlike IFEval which tests familiar constraint types, IFBench specifically measures how well models follow novel instructions they haven't been optimized for, exposing overfitting to common instruction patterns.

### Which model scores highest on IFBench?

MAI-Thinking-1 by Microsoft currently leads with a score of 85% on IFBench.

### How many models are evaluated on IFBench?

40 AI models have been evaluated on IFBench on BenchLM.

### Does IFBench affect BenchLM's overall score?

Yes. IFBench is a weighted benchmark inside the Instruction Following category, which carries 5% of BenchLM's overall score. IFBench itself contributes 70% of that category score.

## Compare Top Models on IFBench

- [MAI-Thinking-1 vs Qwen3.8 Max](/compare/mai-thinking-1-vs-qwen3-8-max)
- [Qwen3.8 Max vs Inkling-Small](/compare/inkling-small-vs-qwen3-8-max)
- [Inkling-Small vs Nemotron 3 Ultra](/compare/inkling-small-vs-nemotron-3-ultra)
- [Nemotron 3 Ultra vs Grok 4.3](/compare/grok-4-3-vs-nemotron-3-ultra)
