Benchmark profile
Instruction-Following Eval (IFEval)
A benchmark of 541 prompts built from 25 verifiable instruction types. It tests whether a model follows checkable constraints such as keyword, length, casing, and response-format requirements.
Data verifiedQwen3.5-27B leads the IFEval leaderboard on BenchLM's July 2026 update with 95%, ahead of Agents-A1 (94.8%) and Qwen3.7 Plus (94.6%), across 24 tracked models.
How to read this leaderboard
Editorial review by Glevd · 2026-07-15
Use IFEval to judge compliance with objectively checkable constraints after matching both scoring choices: prompt versus instruction level, and strict versus loose. An unlabeled IFEval number is incomplete because the four metrics answer different questions and can produce different scores.
Operator receipt: 24 sourced rows are currently displayable on this page; the leading published row is Qwen3.5-27B at 95%.
Honest limit: IFEval covers constraints that can be verified automatically. It does not establish factual accuracy, usefulness, judgment, tone, or multi-turn reliability, and strict checks can reject a substantively compliant answer because of formatting. The current table marks public rows whose prompt/instruction and strict/loose metric view is not recorded; do not treat those values as a precise normalized race.
Top models on IFEval — July 20, 2026
As of July 20, 2026, Qwen3.5-27B leads the IFEval leaderboard with 95% , followed by Agents-A1 (94.8%) and Qwen3.7 Plus (94.6%).
Qwen3.5-27B
Alibaba
IFEval metric view not recorded
qwen3-5-27b
Agents-A1
InternScience
IFEval metric view not recorded
agents-a1
Qwen3.7 Plus
Alibaba
IFEval metric view not recorded
qwen3-7-plus
Leaderboard (24 models)
ScoreAccording to BenchLM.ai, Qwen3.5-27B leads the IFEval benchmark with a score of 95%, followed by Agents-A1 (94.8%) and Qwen3.7 Plus (94.6%). The top models are clustered within 0.4 points, suggesting this benchmark is nearing saturation for frontier models.
24 models have been evaluated on IFEval. The benchmark falls in the Instruction Following category. This category carries a 5% weight in BenchLM.ai's overall scoring system. Within that category, IFEval contributes 35% of the category score, so strong performance here directly affects a model's overall ranking.
About IFEval
Year
2023
Tasks
541 prompts across 25 instruction types
Format
Constrained generation
Difficulty
Instruction precision
IFEval applies deterministic checkers and reports four views of the same responses: prompt-level strict, instruction-level strict, prompt-level loose, and instruction-level loose. Prompt-level scoring requires every instruction in a prompt to pass; loose scoring applies limited normalization before checking.
BenchLM freshness & provenance
Version
IFEval 2023
Refresh cadence
Static
Staleness state
Stale
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does IFEval measure?
IFEval measures whether a model follows checkable constraints in 541 prompts built from 25 instruction types. Examples include using required keywords, avoiding specified terms, meeting length limits, or producing a requested structure. Automatic checkers score each response, which makes the result reproducible for the constraints they can verify.
Which IFEval score should I compare?
Match all four metric choices before comparing models: prompt-level or instruction-level, and strict or loose. Prompt-level scoring requires every instruction in a prompt to pass. Loose scoring applies limited normalization to reduce formatting-related failures. An unlabeled IFEval number is not enough evidence for a precise comparison.
Does a high IFEval score mean a model is a good assistant?
No. A high IFEval score means the model obeyed more objectively checkable constraints under the selected metric. It does not show that the answer was factual, useful, well judged, stylistically appropriate, or reliable across a long conversation. Pair it with broader instruction tests and a workflow-specific prompt set.
Compare Top Models on IFEval
Choose a model with this week’s evidence
Join 2,000+ readers for ranking moves, pricing changes, and the claims that still need proof.
One email each week. Unsubscribe anytime.