Skip to main content
Radar

Every change to the models you run, with its source and its date. Releases, price changes, retirements, API changes, and incidents.Every change to the models you run, with its source.

Follow model changes

Instruction-Following Eval (IFEval)

A benchmark of 541 prompts built from 25 verifiable instruction types. It tests whether a model follows checkable constraints such as keyword, length, casing, and response-format requirements.

Data verified 33 confirmed releases in the last 30 daysSee provider release alerts

The public IFEval snapshot ranks Qwen3.5-27B first at 95%, ahead of Agents-A1 (94.8%) and Agents-A1-4B (94.8%) among 32 models. We mirror the table as display-only evidence; it does not affect overall rankings.

How to read this leaderboard

Editorial review by Glevd · 2026-07-15

Use IFEval to judge compliance with objectively checkable constraints after matching both scoring choices: prompt versus instruction level, and strict versus loose. An unlabeled IFEval number is incomplete because the four metrics answer different questions and can produce different scores.

Operator receipt: 32 sourced rows are currently displayable on this page; the leading published row is Qwen3.5-27B at 95%.

Honest limit: IFEval covers constraints that can be verified automatically. It does not establish factual accuracy, usefulness, judgment, tone, or multi-turn reliability, and strict checks can reject a substantively compliant answer because of formatting. The current table marks public rows whose prompt/instruction and strict/loose metric view is not recorded; do not treat those values as a precise normalized race.

Benchmark score on IFEval — September 15, 2026

We mirror the published score view for IFEval. Qwen3.5-27B leads the public snapshot at 95%, followed by Agents-A1 (94.8%) and Agents-A1-4B (94.8%). We do not use these results to rank models overall.

32 modelsInstruction FollowingStaleDisplay onlyUpdated September 15, 2026

Benchmark score table (32 models)

Score
1
Qwen3.5-27BAlibaba · Open weightIFEval metric view not recorded
95%
2
Agents-A1InternScience · Open weightIFEval metric view not recorded
94.8%
3
Agents-A1-4BInternScience · Open weightIFEval metric view not recorded
94.8%
4
Qwen3.7 PlusAlibaba · ClosedIFEval metric view not recorded
94.6%
5
Qwen3.7 MaxAlibaba · ClosedIFEval metric view not recorded
94.3%
6
Qwen3.6 PlusAlibaba · ClosedIFEval metric view not recorded
94.3%
7
Kimi K2.5Moonshot AI · Open weightIFEval metric view not recorded
93.9%
8
dots3-note PreviewDots Studio · Open weightIFEval metric view not recorded
93.9%
9
o3-miniOpenAI · ClosedIFEval metric view not recorded
93.9%
10
Qwen3.5-122B-A10BAlibaba · Open weightIFEval metric view not recorded
93.4%
11
GLM-5Z.AI · Open weightIFEval metric view not recorded
92.6%
12
Qwen3.5 397BAlibaba · Open weightIFEval metric view not recorded
92.6%
13
K-EXAONE 2.0LG AI Research · Open weightIFEval metric view not recorded
92.4%
14
o1OpenAI · ClosedIFEval metric view not recorded
92.2%
15
Qwen3.5-35B-A3BAlibaba · Open weightIFEval metric view not recorded
91.9%
16
LFM2.5-8B-A1BLiquidAI · Open weightIFEval metric view not recorded
91.8%
17
Claude Opus 4.5Anthropic · ClosedIFEval metric view not recorded
90.9%
18
GPT-4.1 miniOpenAI · ClosedIFEval metric view not recorded
88.5%
19
GPT-4.1OpenAI · ClosedIFEval metric view not recorded
87.4%
20
MiniCPM5-2BOpenBMB · Open weightIFEval metric view not recorded
86.7%
21
DeepSeek V3DeepSeek · Open weightIFEval metric view not recorded
86.1%
22
ZAYA1-8BZyphra · Open weightIFEval metric view not recorded
85.6%
23
GPT-4.1 nanoOpenAI · ClosedIFEval metric view not recorded
83.2%
24
LFM2.5-VL-3BLiquidAI · Open weightIFEval metric view not recorded
82.3%
25
Kanana-2 3B InstructKakao · Open weightIFEval metric view not recorded
81.0%
26
Celeris-1Celeris · ClosedIFEval metric view not recorded
80.8%
27
MiniCPM5-1BOpenBMB · Open weightIFEval metric view not recorded
80.4%
28
Kanana-2 1.3B InstructKakao · Open weightIFEval metric view not recorded
77.6%
29
Mellum2-12B-A2.5B-ThinkingJetBrains · Open weightIFEval metric view not recorded
76.5%
30
Mellum2-12B-A2.5B-InstructJetBrains · Open weightIFEval metric view not recorded
75.8%
31
LFM2.5-230MLiquidAI · Open weightIFEval metric view not recorded
71.7%
32
LFM2.5-VL-450MLiquidAI · Open weightIFEval metric view not recorded
61.2%

The published IFEval snapshot places Qwen3.5-27B first at 95%. The third row is 0.2 points behind. The broader top-10 range is 1.6 points, so many of the published results sit in a relatively narrow band.

32 models have been evaluated on IFEval. The benchmark falls in the Instruction Following category. This category carries a 5% weight in BenchLM.ai's overall scoring system. IFEval is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.

About IFEval

Year

2023

Tasks

541 prompts across 25 instruction types

Format

Constrained generation

Difficulty

Instruction precision

IFEval applies deterministic checkers and reports four views of the same responses: prompt-level strict, instruction-level strict, prompt-level loose, and instruction-level loose. Prompt-level scoring requires every instruction in a prompt to pass; loose scoring applies limited normalization before checking.

BenchLM freshness & provenance

Version

IFEval 2023

Refresh cadence

Static

Staleness state

Stale

Question availability

Public benchmark set

StaleDisplay only

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does IFEval measure?

IFEval measures whether a model follows checkable constraints in 541 prompts built from 25 instruction types. Examples include using required keywords, avoiding specified terms, meeting length limits, or producing a requested structure. Automatic checkers score each response, which makes the result reproducible for the constraints they can verify.

Which IFEval score should I compare?

Match all four metric choices before comparing models: prompt-level or instruction-level, and strict or loose. Prompt-level scoring requires every instruction in a prompt to pass. Loose scoring applies limited normalization to reduce formatting-related failures. An unlabeled IFEval number is not enough evidence for a precise comparison.

Does a high IFEval score mean a model is a good assistant?

No. A high IFEval score means the model obeyed more objectively checkable constraints under the selected metric. It does not show that the answer was factual, useful, well judged, stylistically appropriate, or reliable across a long conversation. Pair it with broader instruction tests and a workflow-specific prompt set.

Last updated: September 15, 2026 · BenchLM version IFEval 2023

Know when it’s worth switching models

The model to choose, the cheaper alternative, and the release we would wait on.

Read a sample issue

Join 2,000+ readers.

One email each week. Unsubscribe anytime.