modelgrep

Best LLMs for Instruction Following — Qwen

Match · Updated August 2026

The best Qwen model for instruction following is Qwen3.7 Max, satisfying 80.5% of IFBench's verifiable prompt constraints. Qwen3.5 397B A17B (78.8%) and Qwen3.7 Plus (78.0%) round out the top three.

80.5%IFBench
46.0Intelligence
207 t/sSpeed
$1.48Input /M
1MContext
  1. 1qwen logo
    qwen3.7-max
    ReasoningToolsJSON46.0 intel · $1.48/M · 207 t/s
    80.5%
    IFBench
  2. 2qwen logo
    qwen3.5-397b-a17b
    ReasoningToolsJSON+133.7 intel · $0.390/M · 69 t/s
    78.8%
    IFBench
  3. 3qwen logo
    qwen3.7-plus
    ReasoningToolsJSON+139.0 intel · $0.320/M · 12 t/s
    78.0%
    IFBench
  4. 4qwen logo
    qwen3.6-max-preview
    ReasoningToolsJSON40.0 intel · $1.03/M · 262K ctx
    76.6%
    IFBench
  5. 5qwen logo
    qwen3.5-122b-a10b
    ReasoningToolsJSON+132.3 intel · $0.260/M · 127 t/s
    75.7%
    IFBench
  6. 6qwen logo
    qwen3.5-27b
    ReasoningToolsJSON+133.8 intel · $0.195/M · 262K ctx
    75.6%
    IFBench
  7. 7qwen logo
    qwen3.6-plus
    ReasoningToolsJSON+139.6 intel · $0.325/M · 55 t/s
    75.2%
    IFBench
  8. 8qwen logo
    qwen3.5-35b-a3b
    ReasoningToolsJSON+129.3 intel · $0.140/M · 262K ctx
    72.5%
    IFBench
  9. 9qwen logo
    qwen3-max-thinking
    ReasoningToolsJSON31.7 intel · $0.780/M · 262K ctx
    70.7%
    IFBench
  10. 10qwen logo
    qwen3.6-27b
    ReasoningToolsJSON+137.1 intel · $0.289/M · 60 t/s
    67.6%
    IFBench
  11. 11qwen logo
    qwen3.5-9b
    ReasoningToolsJSON+121.4 intel · $0.100/M · 74 t/s
    66.7%
    IFBench
  12. 12qwen logo
    qwen3.6-35b-a3b
    ReasoningToolsJSON+131.6 intel · $0.140/M · 158 t/s
    64.4%
    IFBench
  13. 13qwen logo
    qwen3-max
    ToolsJSON24.0 intel · $0.780/M · 29 t/s
    44.1%
    IFBench
  14. 14qwen logo
    qwen3-vl-235b-a22b-instruct
    ToolsJSONVision14.3 intel · $0.210/M · 19 t/s
    42.7%
    IFBench
  15. 15qwen logo
    qwen3-next-80b-a3b-instruct
    ToolsJSON13.7 intel · $0.090/M · 74 t/s
    39.7%
    IFBench
  16. 16qwen logo
    qwen3-vl-32b-instruct
    ToolsJSONVision11.1 intel · $0.104/M · 38 t/s
    39.2%
    IFBench
  17. 17qwen logo
    qwen-2.5-72b-instruct
    ToolsJSON9.6 intel · $0.360/M · 33K ctx
    36.9%
    IFBench
  18. 18qwen logo
    qwen3-coder-next
    ToolsJSON21.1 intel · $0.120/M · 133 t/s
    35.2%
    IFBench
  19. 19qwen logo
    qwen3-vl-30b-a3b-instruct
    ToolsJSONVision10.0 intel · $0.150/M · 34 t/s
    33.1%
    IFBench
  20. 20qwen logo
    qwen3-coder-30b-a3b-instruct
    ToolsJSON13.6 intel · $0.070/M · 81 t/s
    32.7%
    IFBench
  21. 21qwen logo
    qwen3-vl-8b-instruct
    ToolsJSONVision8.4 intel · $0.117/M · 42 t/s
    32.3%
    IFBench

How this is ranked

AI models ranked by IFBench — how reliably a model obeys explicit, verifiable constraints in the prompt (format, length, inclusion and exclusion rules). High scores mean fewer retries and less prompt-wrangling in production, which often matters more than raw intelligence for pipelines.

Frequently asked

Which Qwen model follows instructions most reliably?

The best Qwen model for instruction following is Qwen3.7 Max, satisfying 80.5% of IFBench's verifiable prompt constraints. Qwen3.5 397B A17B (78.8%) and Qwen3.7 Plus (78.0%) round out the top three.

What's a good alternative to Qwen3.7 Max?

Qwen3.5 397B A17B (78.8%) is the closest alternative on this metric, followed by Qwen3.7 Plus (78.0%). See the full ranking above for the tradeoffs.

How many Qwen models are there?

modelgrep tracks 49 Qwen models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Qwen3.7 Max. 21 of them qualify for this ranking.

More Qwen rankings

All rankings