modelgrep

Best LLMs for Instruction Following

Match · Updated August 2026

The best LLM for instruction following is MiniMax M3, satisfying 82.9% of IFBench's verifiable prompt constraints. Grok 4.3 (81.3%) and Grok 4.20 (81.2%) round out the top three.

82.9%IFBench
44.4Intelligence
101 t/sSpeed
$0.300Input /M
1.0MContext
  1. 1minimax logo
    minimax-m3
    ReasoningToolsJSON+144.4 intel · $0.300/M · 101 t/s
    82.9%
    IFBench
  2. 2x-ai logo
    grok-4.3
    ReasoningToolsJSON+137.6 intel · $1.25/M · 1M ctx
    81.3%
    IFBench
  3. 3x-ai logo
    grok-4.20
    ReasoningToolsJSON+137.0 intel · $1.25/M · 2M ctx
    81.2%
    IFBench
  4. 4qwen logo
    qwen3.7-max
    ReasoningToolsJSON46.0 intel · $1.48/M · 207 t/s
    80.5%
    IFBench
  5. 5xiaomi logo
    mimo-v2.5-pro
    ReasoningToolsJSON42.2 intel · $0.435/M · 65 t/s
    79.9%
    IFBench
  6. 6qwen logo
    qwen3.5-397b-a17b
    ReasoningToolsJSON+133.7 intel · $0.390/M · 69 t/s
    78.8%
    IFBench
  7. 7qwen logo
    qwen3.7-plus
    ReasoningToolsJSON+139.0 intel · $0.320/M · 12 t/s
    78.0%
    IFBench
  8. 8openai logo
    gpt-5.2-codex
    ReasoningToolsJSON+140.1 intel · $1.75/M · 400K ctx
    77.6%
    IFBench
  9. 9google logo
    gemini-3.1-flash-lite
    ReasoningToolsJSON+225.0 intel · $0.250/M · 337 t/s
    77.2%
    IFBench
  10. 10google logo
    gemini-3.1-flash-lite-preview
    ReasoningToolsJSON+225.0 intel · $0.250/M · 337 t/s
    77.2%
    IFBench
  11. 11google logo
    gemini-3.1-pro-preview
    ReasoningToolsJSON+246.5 intel · $2.00/M · 136 t/s
    77.1%
    IFBench
  12. 12qwen logo
    qwen3.6-max-preview
    ReasoningToolsJSON40.0 intel · $1.03/M · 262K ctx
    76.6%
    IFBench
  13. 13deepseek logo
    deepseek-v4-pro
    ReasoningToolsJSON44.3 intel · $0.435/M · 67 t/s
    76.5%
    IFBench
  14. 14google logo
    gemini-3.5-flash
    ReasoningToolsJSON+250.2 intel · $1.50/M · 275 t/s
    76.3%
    IFBench
  15. 15z-ai logo
    glm-5.1
    ReasoningToolsJSON40.2 intel · $0.966/M · 205K ctx
    76.3%
    IFBench
  16. 16moonshotai logo
    kimi-k2.6
    ReasoningToolsJSON+144.2 intel · $0.589/M · 262K ctx
    76.0%
    IFBench
  17. 17openai logo
    gpt-5.4-nano
    ReasoningToolsJSON+138.2 intel · $0.200/M · 400K ctx
    75.9%
    IFBench
  18. 18openai logo
    gpt-5.5
    ReasoningToolsJSON+154.8 intel · $5.00/M · 1.1M ctx
    75.9%
    IFBench
  19. 19minimax logo
    minimax-m2.7
    ReasoningToolsJSON38.1 intel · $0.270/M · 205K ctx
    75.7%
    IFBench
  20. 20qwen logo
    qwen3.5-122b-a10b
    ReasoningToolsJSON+132.3 intel · $0.260/M · 127 t/s
    75.7%
    IFBench
  21. 21qwen logo
    qwen3.5-27b
    ReasoningToolsJSON+133.8 intel · $0.195/M · 262K ctx
    75.6%
    IFBench
  22. 22openai logo
    gpt-5.2
    ReasoningToolsJSON+142.2 intel · $1.75/M · 74 t/s
    75.4%
    IFBench
  23. 23openai logo
    gpt-5-mini
    ReasoningToolsJSON+125.3 intel · $0.250/M · 84 t/s
    75.4%
    IFBench
  24. 24openai logo
    gpt-5.3-codex
    ReasoningToolsJSON+144.3 intel · $1.75/M · 122 t/s
    75.4%
    IFBench
  25. 25qwen logo
    qwen3.6-plus
    ReasoningToolsJSON+139.6 intel · $0.325/M · 55 t/s
    75.2%
    IFBench

How this is ranked

AI models ranked by IFBench — how reliably a model obeys explicit, verifiable constraints in the prompt (format, length, inclusion and exclusion rules). High scores mean fewer retries and less prompt-wrangling in production, which often matters more than raw intelligence for pipelines.

Frequently asked

Which LLM follows instructions most reliably?

The best LLM for instruction following is MiniMax M3, satisfying 82.9% of IFBench's verifiable prompt constraints. Grok 4.3 (81.3%) and Grok 4.20 (81.2%) round out the top three.

Which AI model follows instructions best?

The best AI model for instruction following is MiniMax M3, satisfying 82.9% of IFBench's verifiable prompt constraints. Grok 4.3 (81.3%) and Grok 4.20 (81.2%) round out the top three.

What's a good alternative to MiniMax M3?

Grok 4.3 (81.3%) is the closest alternative on this metric, followed by Grok 4.20 (81.2%). See the full ranking above for the tradeoffs.

By maker

All rankings