modelgrep

Best LLMs for Instruction Following — OpenAI

Match · Updated August 2026

The best OpenAI model for instruction following is GPT-5.2-Codex, satisfying 77.6% of IFBench's verifiable prompt constraints. GPT-5.4 Nano (75.9%) and GPT-5.5 (75.9%) round out the top three.

77.6%IFBench
40.1Intelligence
$1.75Input /M
400KContext
  1. 1openai logo
    gpt-5.2-codex
    ReasoningToolsJSON+140.1 intel · $1.75/M · 400K ctx
    77.6%
    IFBench
  2. 2openai logo
    gpt-5.4-nano
    ReasoningToolsJSON+138.2 intel · $0.200/M · 400K ctx
    75.9%
    IFBench
  3. 3openai logo
    gpt-5.5
    ReasoningToolsJSON+154.8 intel · $5.00/M · 1.1M ctx
    75.9%
    IFBench
  4. 4openai logo
    gpt-5.2
    ReasoningToolsJSON+142.2 intel · $1.75/M · 74 t/s
    75.4%
    IFBench
  5. 5openai logo
    gpt-5-mini
    ReasoningToolsJSON+125.3 intel · $0.250/M · 84 t/s
    75.4%
    IFBench
  6. 6openai logo
    gpt-5.3-codex
    ReasoningToolsJSON+144.3 intel · $1.75/M · 122 t/s
    75.4%
    IFBench
  7. 7openai logo
    gpt-5.4
    ReasoningToolsJSON+151.4 intel · $2.50/M · 1.1M ctx
    73.9%
    IFBench
  8. 8openai logo
    gpt-5.4-mini
    ReasoningToolsJSON+140.0 intel · $0.750/M · 400K ctx
    73.3%
    IFBench
  9. 9openai logo
    gpt-5
    ReasoningToolsJSON+134.7 intel · $1.25/M · 65 t/s
    73.1%
    IFBench
  10. 10openai logo
    gpt-5.1
    ReasoningToolsJSON+136.9 intel · $1.25/M · 51 t/s
    72.9%
    IFBench
  11. 11openai logo
    gpt-5.6-sol
    ReasoningToolsJSON+158.9 intel · $5.00/M · 50 t/s
    72.7%
    IFBench
  12. 12openai logo
    o3
    ReasoningToolsJSON+130.4 intel · $2.00/M · 44 t/s
    71.4%
    IFBench
  13. 13openai logo
    gpt-5.6-terra
    ReasoningToolsJSON+155.0 intel · $1.00/M · 60 t/s
    71.2%
    IFBench
  14. 14openai logo
    o1
    ReasoningToolsJSON+123.4 intel · $15.00/M · 200K ctx
    70.3%
    IFBench
  15. 15openai logo
    gpt-5.1-codex
    ReasoningToolsJSON+134.7 intel · $1.25/M · 66 t/s
    70.0%
    IFBench
  16. 16openai logo
    gpt-oss-120b
    ReasoningToolsJSON23.8 intel · $0.037/M · 490 t/s
    69.0%
    IFBench
  17. 17openai logo
    o4-mini-high
    ReasoningToolsJSON+125.6 intel · $1.10/M · 95 t/s
    68.7%
    IFBench
  18. 18openai logo
    o4-mini
    ReasoningToolsJSON+125.6 intel · $1.10/M · 81 t/s
    68.7%
    IFBench
  19. 19openai logo
    gpt-5.1-codex-mini
    ReasoningToolsJSON+130.6 intel · $0.250/M · 51 t/s
    67.9%
    IFBench
  20. 20openai logo
    gpt-5-nano
    ReasoningToolsJSON+119.9 intel · $0.050/M · 81 t/s
    67.6%
    IFBench
  21. 21openai logo
    o3-mini-high
    ReasoningToolsJSON15.6 intel · $1.10/M · 200K ctx
    67.1%
    IFBench
  22. 22openai logo
    gpt-oss-20b
    ReasoningToolsJSON14.9 intel · $0.030/M · 274 t/s
    65.1%
    IFBench
  23. 23openai logo
    gpt-oss-20b:free
    ReasoningToolsJSON14.9 intel · Free/M · 17 t/s
    65.1%
    IFBench
  24. 24openai logo
    gpt-4.1
    ToolsJSONVision19.4 intel · $2.00/M · 1.0M ctx
    43.0%
    IFBench
  25. 25openai logo
    gpt-4.1-mini
    ToolsJSONVision14.8 intel · $0.400/M · 1.0M ctx
    38.3%
    IFBench

How this is ranked

AI models ranked by IFBench — how reliably a model obeys explicit, verifiable constraints in the prompt (format, length, inclusion and exclusion rules). High scores mean fewer retries and less prompt-wrangling in production, which often matters more than raw intelligence for pipelines.

Frequently asked

Which OpenAI model follows instructions most reliably?

The best OpenAI model for instruction following is GPT-5.2-Codex, satisfying 77.6% of IFBench's verifiable prompt constraints. GPT-5.4 Nano (75.9%) and GPT-5.5 (75.9%) round out the top three.

What's a good alternative to GPT-5.2-Codex?

GPT-5.4 Nano (75.9%) is the closest alternative on this metric, followed by GPT-5.5 (75.9%). See the full ranking above for the tradeoffs.

How many OpenAI models are there?

modelgrep tracks 60 OpenAI models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by GPT-5.6 Sol. 25 of them qualify for this ranking.

More OpenAI rankings

All rankings