Best LLMs for Instruction Following — September 2026 Leaderboard
Data refreshed:
Ability to follow precise instructions and constraints
As of September 2026, the top instruction following model on the BenchLM leaderboard is MAI-Thinking-1 with a weighted instruction following score of 94.7.
Decision lens: use provisional-ranked mode for broader public evidence and verified-ranked mode for source-only comparisons. A model can move between views as evidence coverage changes.
- Data refreshed
- September 4, 2026
- Provisional-ranked
- 120 of 417 models
- Verified-ranked
- 124 of 417 models
- Weighted evidence
- 1 of 3 benchmarks
3 tracked benchmarks
IFEval, IFBench, SOB Value Acc
Evidence set: IFEval, IFBench, SOB Value Acc
Best Instruction Following picks
BenchLM summaries for instruction following plus the practical tradeoffs users check next: open weights, price, speed, latency, and context.
IFEval Leaderboard
Primary score: weighted instruction following score. Higher values rank first. Use the Show metric control to change the value shown in each row.
Switch between provisional-ranked and verified-ranked modes to compare the broader public dataset with sourced-only rankings.
Filters
1 MAI-Thinking-1 Microsoft | 94.7% | 51 | — | 85% | — |
|---|---|---|---|---|---|
2 Grok 4.3 xAI | 94.4% | 63 | — | 81.3% | — |
3 GPT-5.2-Codex OpenAI | 93.5% | 70 | 92%P | — | — |
4 | 93.5% | 61 | — | — | — |
| 93.5% | 60 | 92.6%P | — | — | |
6 MiMo-V2.5-Pro Xiaomi | 93.5% | 62 | — | — | — |
7 GPT-5.5 OpenAI | 92.9% | 69 | — | — | — |
8 Muse Spark Meta | 92.9% | 55 | — | — | — |
9 GPT-5.4 nano OpenAI | 92.9% | 49 | — | — | — |
10 | 92.7% | 56 | — | 75.7%P | — |
11 | 92.7% | 49 | 93.4% | — | — |
12 | 92.6% | 51 | — | — | — |
13 | 92.6% | 50 | 95% | — | — |
14 GPT-5.2 OpenAI | 92.3% | 51 | 94%P | 76.2%P | — |
15 GPT-5.3 Codex OpenAI | 92.3% | 68 | 93%P | 78.1%P | — |
16 Qwen3.7 Max Alibaba | 91.1% | 72 | 94.3% | 79.1% | — |
17 Qwen3.7 Plus Alibaba | 91.1% | 67 | 94.6% | 79.1% | — |
18 | 90.7% | 77 | — | 82.8% | — |
19 GPT-5.4 OpenAI | 90.4% | 63 | 96%P | 79.4%P | — |
20 | 90.4% | 46 | — | — | — |
21 | 89.8% | 47 | — | — | — |
| 89.6% | 67 | — | — | — | |
23 GPT-5.4 mini OpenAI | 89.6% | 60 | 87.4%P | — | — |
24 | 89.6% | 56 | — | 82.2% | — |
25 GLM-5-Turbo Z.AI | 89.4% | 63 | — | — | — |
Top AI Models for Instruction Following — September 2026
As of September 2026, MAI-Thinking-1 leads the provisional instruction following leaderboard with a score of 94.7%, followed by Grok 4.3 (94.4%) and GPT-5.2-Codex (93.5%). BenchLM is currently showing 120 provisional-ranked models and 124 verified-ranked models in this category.
Strong instruction-following accuracy on the current weighted evidence.
What changed
MAI-Thinking-1 leads the current instruction-following table on its sourced weighted row.
GPT-5.4 close second on IFEval and IFBench.
Claude Opus 4.6 holds #3, with strong IFBench scores.
How to choose
Top models by benchmark
Instruction-following benchmark used in first-party comparison charts for agent-oriented reasoning models.(70% of category score)
Score in Context
What these scores mean
Instruction following carries a 5% weight in overall scoring — small, but it directly measures reliability. A model that ignores formatting rules, word count limits, or inclusion constraints is unusable in automated pipelines. The weighted score blends IFEval (verifiable constraints) and IFBench.
Known limitations
IFEval tests a specific set of verifiable constraints (word count, formatting, inclusion/exclusion) — it doesn't capture the full range of instruction-following quality. A model can score well on IFEval but still misinterpret nuanced or ambiguous instructions. IFBench is newer and coverage is still building.
How we weight
Instruction following carries a 5% weight in BenchLM.ai's overall scoring. While a smaller category, it directly measures reliability — a model that ignores constraints is unusable in production pipelines. See the instruction following leaderboard.
Leaderboards exclude benchmark rows that BenchLM generated from other scores or cloned from reference models. When a weighted benchmark is missing after that filter, the category falls back to the remaining trustworthy public rows instead of filling the gap with synthetic values.
The full scoring rules, freshness handling, and runtime/pricing caveats live on the BenchLM methodology page.
Scroll horizontally to read the full evidence ledger.
| Benchmark | Weight | Status | Description |
|---|---|---|---|
| IFEval | — | Display only | Tests ability to follow verifiable instructions like format constraints and content requirements |
| IFBench | 70% | Weighted | Instruction-following benchmark used in first-party comparison charts for agent-oriented reasoning models. |
| SOB Value Acc | — | Display only | A structured-output benchmark from Interfaze measuring whether extracted JSON leaf values exactly match verified ground truth. |
About Instruction Following Benchmarks
Tests ability to follow verifiable instructions like format constraints and content requirements
Common questions
What is the best LLM for instruction following?
The best instruction-following LLMs are ranked by benchmarks like IFEval, which measure how accurately models adhere to specific formatting, length, and content constraints.
What is IFEval and how does it work?
IFEval (Instruction Following Evaluation) tests whether LLMs can precisely follow verifiable instructions such as word count limits, formatting rules, and inclusion/exclusion constraints.
Why is instruction following important for LLMs?
Instruction following is critical because real-world applications require models to reliably adhere to user specifications, output formats, and constraints rather than just generating plausible text.
Instruction following updates
Instruction following scores just got interesting. Get the weekly update.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.