Benchmark Guide
What Is IFEval? How LLM Instruction Following Is Measured
Last updated: October 2, 2026 β model data is refreshed automatically from the id8 dataset
A model that writes well but ignores 'answer in three bullet points' is hard to build a product on. IFEval measures exactly that: whether the model does what the instruction says. It is one of the most practical benchmarks for anyone writing prompts for an application.
Models ranked on Arena Elo, IFEval and SimpleQA Verified.
See the chat ranking βThe idea
Most ways of judging writing need a human or another model as judge, which is slow and subjective. IFEval avoids that by using only instructions whose result a simple program can check.
Examples of such instructions:
- Write at least 300 words.
- Use exactly three bullet points.
- Include the words "budget" and "deadline".
- Do not use any commas.
- Answer in all lowercase letters.
- Wrap the whole answer in JSON.
- End with the sentence "Is there anything else I can help with?"
- Answer in Hindi.
The benchmark has 25 types of instruction like these and about 500 prompts. Many prompts combine several instructions, which is where models start to fail.
How it is scored
For each prompt, a checker tests every instruction.
- Prompt-level accuracy: the share of prompts where all instructions were followed.
- Instruction-level accuracy: the share of individual instructions followed.
Each is reported in a strict form and a loose form. Loose scoring strips common extras, such as a leading "Sure, here isβ¦" line, before checking. The single figure usually quoted is strict prompt-level accuracy.
Current top scores
| # | Model | Provider | IFEval |
|---|---|---|---|
| 1 | Claude Opus 4.6 | Anthropic | 94 |
| 2 | Kimi K2.5 | Moonshot AI | 94 |
| 3 | Qwen3.5 122B-A10B | Alibaba | 93.4 |
| 4 | Qwen 3.5 | Alibaba | 92.6 |
| 5 | Llama 3.3 70B | Meta | 92.1 |
| 6 | Qwen3.5 9B | Alibaba | 91.5 |
| 7 | Gemma 3 4B | 90.2 | |
| 8 | Nemotron Ultra 253B | NVIDIA | 89.5 |
| 9 | Gemma 3 12B | 88.9 | |
| 10 | GLM-5 | Zhipu AI | 88 |
27 of the 164 models we track have a published IFEval score. The newest flagship models often do not report it.
Why it matters for real products
If you build with an LLM, your code depends on the shape of the output:
- A parser expects JSON and receives a sentence before it.
- A user interface has room for 50 words and receives 200.
- A workflow expects one of four labels and receives a fifth.
Each of these is an instruction-following failure, and IFEval is the closest public measure of how often a model commits them. It is particularly relevant for retrieval systems and agents, where instructions such as "answer only from the passages" or "call one tool at a time" must hold on every request.
Limits
- Format, not content. IFEval cannot tell whether the 300 words are true or useful.
- Simple instructions. Real system prompts contain subtler rules, such as tone or what to refuse, which a program cannot check.
- Near the ceiling. Leading models score high enough that small differences are not meaningful.
- English-centred. Most prompts are in English, so the score says little about following instructions in Hindi or other languages.
- Public since 2023. Models may have seen the prompts in training.
Getting better instruction following from any model
- Put format rules at the end of the prompt, close to where the answer begins.
- Give one short example of the output you want.
- Use the provider's structured-output or JSON mode where it exists, which enforces format rather than requesting it.
- Keep the number of simultaneous constraints small.
- Validate the output in code and retry on failure.
How id8 uses it
IFEval contributes to our grades for chat, writing, summarisation and tool calling. Where a model has no published score, the ranking estimates from its other results. See Best LLM for Chat for the table.
Frequently asked questions
Which model scores highest on IFEval?
As of October 1, 2026, the highest published IFEval score in the id8 dataset is Claude Opus 4.6 (94).
What does IFEval stand for?
Instruction-Following Evaluation. It was published by researchers at Google in 2023.
How many prompts does IFEval have?
About 500 prompts, built from 25 kinds of instruction that a program can verify.
Does IFEval check whether the answer is good?
No. It checks only whether the stated instructions were obeyed. An answer can follow every instruction and still be wrong or badly written.
What is the difference between strict and loose accuracy?
Strict accuracy checks the response exactly as written. Loose accuracy first removes common wrappers, such as an opening line or markdown markers, so the model is not failed for a harmless preamble.