Instruction-Following Eval (IFEval)
A benchmark of 541 prompts built from 25 verifiable instruction types. It tests whether a model follows checkable constraints such as keyword, length, casing, and response-format requirements.
Data verified 33 confirmed releases in the last 30 daysSee provider release alertsThe public IFEval snapshot ranks Qwen3.5-27B first at 95%, ahead of Agents-A1 (94.8%) and Agents-A1-4B (94.8%) among 32 models. We mirror the table as display-only evidence; it does not affect overall rankings.
How to read this leaderboard
Editorial review by Glevd · 2026-07-15
Use IFEval to judge compliance with objectively checkable constraints after matching both scoring choices: prompt versus instruction level, and strict versus loose. An unlabeled IFEval number is incomplete because the four metrics answer different questions and can produce different scores.
Operator receipt: 32 sourced rows are currently displayable on this page; the leading published row is Qwen3.5-27B at 95%.
Honest limit: IFEval covers constraints that can be verified automatically. It does not establish factual accuracy, usefulness, judgment, tone, or multi-turn reliability, and strict checks can reject a substantively compliant answer because of formatting. The current table marks public rows whose prompt/instruction and strict/loose metric view is not recorded; do not treat those values as a precise normalized race.
Benchmark score on IFEval — September 15, 2026
We mirror the published score view for IFEval. Qwen3.5-27B leads the public snapshot at 95%, followed by Agents-A1 (94.8%) and Agents-A1-4B (94.8%). We do not use these results to rank models overall.
Qwen3.5-27B
Alibaba
IFEval metric view not recorded
Agents-A1
InternScience
IFEval metric view not recorded
Agents-A1-4B
InternScience
IFEval metric view not recorded
32 modelsInstruction FollowingStaleDisplay onlyUpdated September 15, 2026
Benchmark score table (32 models)
ScoreThe published IFEval snapshot places Qwen3.5-27B first at 95%. The third row is 0.2 points behind. The broader top-10 range is 1.6 points, so many of the published results sit in a relatively narrow band.
32 models have been evaluated on IFEval. The benchmark falls in the Instruction Following category. This category carries a 5% weight in BenchLM.ai's overall scoring system. IFEval is currently displayed for reference but excluded from the scoring formula, so it does not directly affect overall rankings.
About IFEval
Year
2023
Tasks
541 prompts across 25 instruction types
Format
Constrained generation
Difficulty
Instruction precision
IFEval applies deterministic checkers and reports four views of the same responses: prompt-level strict, instruction-level strict, prompt-level loose, and instruction-level loose. Prompt-level scoring requires every instruction in a prompt to pass; loose scoring applies limited normalization before checking.
BenchLM freshness & provenance
Version
IFEval 2023
Refresh cadence
Static
Staleness state
Stale
Question availability
Public benchmark set
BenchLM uses freshness metadata to decide whether a benchmark should still be treated as a strong differentiator, a benchmark to watch, or a display-only reference. For the full scoring policy, see the BenchLM methodology page.
FAQ
What does IFEval measure?
IFEval measures whether a model follows checkable constraints in 541 prompts built from 25 instruction types. Examples include using required keywords, avoiding specified terms, meeting length limits, or producing a requested structure. Automatic checkers score each response, which makes the result reproducible for the constraints they can verify.
Which IFEval score should I compare?
Match all four metric choices before comparing models: prompt-level or instruction-level, and strict or loose. Prompt-level scoring requires every instruction in a prompt to pass. Loose scoring applies limited normalization to reduce formatting-related failures. An unlabeled IFEval number is not enough evidence for a precise comparison.
Does a high IFEval score mean a model is a good assistant?
No. A high IFEval score means the model obeyed more objectively checkable constraints under the selected metric. It does not show that the answer was factual, useful, well judged, stylistically appropriate, or reliable across a long conversation. Pair it with broader instruction tests and a workflow-specific prompt set.
Compare Top Models on IFEval
Know when it’s worth switching models
The model to choose, the cheaper alternative, and the release we would wait on.
Read a sample issueJoin 2,000+ readers.
One email each week. Unsubscribe anytime.