Skip to main content

Benchmark profile

Instruction-Following Eval (IFEval)

A benchmark of 541 prompts built from 25 verifiable instruction types. It tests whether a model follows checkable constraints such as keyword, length, casing, and response-format requirements.

Data verified

Qwen3.5-27B leads the IFEval leaderboard on BenchLM's July 2026 update with 95%, ahead of Agents-A1 (94.8%) and Qwen3.7 Plus (94.6%), across 24 tracked models.

How to read this leaderboard

Editorial review by Glevd · 2026-07-15

Use IFEval to judge compliance with objectively checkable constraints after matching both scoring choices: prompt versus instruction level, and strict versus loose. An unlabeled IFEval number is incomplete because the four metrics answer different questions and can produce different scores.

Operator receipt: 24 sourced rows are currently displayable on this page; the leading published row is Qwen3.5-27B at 95%.

Honest limit: IFEval covers constraints that can be verified automatically. It does not establish factual accuracy, usefulness, judgment, tone, or multi-turn reliability, and strict checks can reject a substantively compliant answer because of formatting. The current table marks public rows whose prompt/instruction and strict/loose metric view is not recorded; do not treat those values as a precise normalized race.

Top models on IFEval — July 20, 2026

As of July 20, 2026, Qwen3.5-27B leads the IFEval leaderboard with 95% , followed by Agents-A1 (94.8%) and Qwen3.7 Plus (94.6%).

24 modelsInstruction Following35% of category scoreStaleUpdated July 20, 2026

Leaderboard (24 models)

Score
1
Qwen3.5-27BAlibaba · Open weightIFEval metric view not recorded
95%
2
Agents-A1InternScience · Open weightIFEval metric view not recorded
94.8%
3
Qwen3.7 PlusAlibaba · ClosedIFEval metric view not recorded
94.6%
4
Qwen3.7 MaxAlibaba · ClosedIFEval metric view not recorded
94.3%
5
Qwen3.6 PlusAlibaba · ClosedIFEval metric view not recorded
94.3%
6
Kimi K2.5Moonshot AI · Open weightIFEval metric view not recorded
93.9%
7
o3-miniOpenAI · ClosedIFEval metric view not recorded
93.9%
8
Qwen3.5-122B-A10BAlibaba · Open weightIFEval metric view not recorded
93.4%
9
GLM-5Z.AI · Open weightIFEval metric view not recorded
92.6%
10
Qwen3.5 397BAlibaba · Open weightIFEval metric view not recorded
92.6%
11
o1OpenAI · ClosedIFEval metric view not recorded
92.2%
12
Qwen3.5-35B-A3BAlibaba · Open weightIFEval metric view not recorded
91.9%
13
LFM2.5-8B-A1BLiquidAI · Open weightIFEval metric view not recorded
91.8%
14
Claude Opus 4.5Anthropic · ClosedIFEval metric view not recorded
90.9%
15
GPT-4.1 miniOpenAI · ClosedIFEval metric view not recorded
88.5%
16
GPT-4.1OpenAI · ClosedIFEval metric view not recorded
87.4%
17
DeepSeek V3DeepSeek · Open weightIFEval metric view not recorded
86.1%
18
ZAYA1-8BZyphra · Open weightIFEval metric view not recorded
85.6%
19
GPT-4.1 nanoOpenAI · ClosedIFEval metric view not recorded
83.2%
20
MiniCPM5-1BOpenBMB · Open weightIFEval metric view not recorded
80.4%
21
Mellum2-12B-A2.5B-ThinkingJetBrains · Open weightIFEval metric view not recorded
76.5%
22
Mellum2-12B-A2.5B-InstructJetBrains · Open weightIFEval metric view not recorded
75.8%
23
LFM2.5-230MLiquidAI · Open weightIFEval metric view not recorded
71.7%
24
LFM2.5-VL-450MLiquidAI · Open weightIFEval metric view not recorded
61.2%

According to BenchLM.ai, Qwen3.5-27B leads the IFEval benchmark with a score of 95%, followed by Agents-A1 (94.8%) and Qwen3.7 Plus (94.6%). The top models are clustered within 0.4 points, suggesting this benchmark is nearing saturation for frontier models.

24 models have been evaluated on IFEval. The benchmark falls in the Instruction Following category. This category carries a 5% weight in BenchLM.ai's overall scoring system. Within that category, IFEval contributes 35% of the category score, so strong performance here directly affects a model's overall ranking.

About IFEval

Year

2023

Tasks

541 prompts across 25 instruction types

Format

Constrained generation

Difficulty

Instruction precision

IFEval applies deterministic checkers and reports four views of the same responses: prompt-level strict, instruction-level strict, prompt-level loose, and instruction-level loose. Prompt-level scoring requires every instruction in a prompt to pass; loose scoring applies limited normalization before checking.

BenchLM freshness & provenance

Version

IFEval 2023

Refresh cadence

Static

Staleness state

Stale

Question availability

Public benchmark set

Stale

BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.

FAQ

What does IFEval measure?

IFEval measures whether a model follows checkable constraints in 541 prompts built from 25 instruction types. Examples include using required keywords, avoiding specified terms, meeting length limits, or producing a requested structure. Automatic checkers score each response, which makes the result reproducible for the constraints they can verify.

Which IFEval score should I compare?

Match all four metric choices before comparing models: prompt-level or instruction-level, and strict or loose. Prompt-level scoring requires every instruction in a prompt to pass. Loose scoring applies limited normalization to reduce formatting-related failures. An unlabeled IFEval number is not enough evidence for a precise comparison.

Does a high IFEval score mean a model is a good assistant?

No. A high IFEval score means the model obeyed more objectively checkable constraints under the selected metric. It does not show that the answer was factual, useful, well judged, stylistically appropriate, or reliable across a long conversation. Pair it with broader instruction tests and a workflow-specific prompt set.

Last updated: July 20, 2026 · BenchLM version IFEval 2023

Choose a model with this week’s evidence

Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.

One email each week. Unsubscribe anytime.