modelgrep

Best LLMs for Tool Calling — Qwen

Match · Updated August 2026

The best Qwen model for tool calling is Qwen3.6 Plus, completing 97.7% of Tau²-Bench's multi-turn tool-use tasks. Qwen3.6 Max Preview (95.9%) and Qwen3.5 397B A17B (95.6%) round out the top three.

97.7%τ²-Bench
39.6Intelligence
55 t/sSpeed
$0.325Input /M
1MContext
  1. 1qwen logo
    qwen3.6-plus
    ReasoningToolsJSON+139.6 intel · $0.325/M · 55 t/s
    97.7%
    τ²-Bench
  2. 2qwen logo
    qwen3.6-max-preview
    ReasoningToolsJSON40.0 intel · $1.03/M · 262K ctx
    95.9%
    τ²-Bench
  3. 3qwen logo
    qwen3.5-397b-a17b
    ReasoningToolsJSON+133.7 intel · $0.390/M · 69 t/s
    95.6%
    τ²-Bench
  4. 4qwen logo
    qwen3.6-35b-a3b
    ReasoningToolsJSON+131.6 intel · $0.140/M · 158 t/s
    95.3%
    τ²-Bench
  5. 5qwen logo
    qwen3.7-max
    ReasoningToolsJSON46.0 intel · $1.48/M · 207 t/s
    94.7%
    τ²-Bench
  6. 6qwen logo
    qwen3.6-27b
    ReasoningToolsJSON+137.1 intel · $0.289/M · 60 t/s
    94.2%
    τ²-Bench
  7. 7qwen logo
    qwen3.5-27b
    ReasoningToolsJSON+133.8 intel · $0.195/M · 262K ctx
    93.9%
    τ²-Bench
  8. 8qwen logo
    qwen3.5-122b-a10b
    ReasoningToolsJSON+132.3 intel · $0.260/M · 127 t/s
    93.6%
    τ²-Bench
  9. 9qwen logo
    qwen3.7-plus
    ReasoningToolsJSON+139.0 intel · $0.320/M · 12 t/s
    93.0%
    τ²-Bench
  10. 10qwen logo
    qwen3.5-35b-a3b
    ReasoningToolsJSON+129.3 intel · $0.140/M · 262K ctx
    89.2%
    τ²-Bench
  11. 11qwen logo
    qwen3.5-9b
    ReasoningToolsJSON+121.4 intel · $0.100/M · 74 t/s
    86.8%
    τ²-Bench
  12. 12qwen logo
    qwen3-max-thinking
    ReasoningToolsJSON31.7 intel · $0.780/M · 262K ctx
    83.6%
    τ²-Bench
  13. 13qwen logo
    qwen3-coder-next
    ToolsJSON21.1 intel · $0.120/M · 133 t/s
    79.5%
    τ²-Bench
  14. 14qwen logo
    qwen3-max
    ToolsJSON24.0 intel · $0.780/M · 29 t/s
    74.3%
    τ²-Bench
  15. 15qwen logo
    qwen3-vl-235b-a22b-instruct
    ToolsJSONVision14.3 intel · $0.210/M · 19 t/s
    35.1%
    τ²-Bench
  16. 16qwen logo
    qwen3-coder-30b-a3b-instruct
    ToolsJSON13.6 intel · $0.070/M · 81 t/s
    34.5%
    τ²-Bench
  17. 17qwen logo
    qwen-2.5-72b-instruct
    ToolsJSON9.6 intel · $0.360/M · 33K ctx
    34.5%
    τ²-Bench
  18. 18qwen logo
    qwen3-vl-32b-instruct
    ToolsJSONVision11.1 intel · $0.104/M · 38 t/s
    29.2%
    τ²-Bench
  19. 19qwen logo
    qwen3-vl-8b-instruct
    ToolsJSONVision8.4 intel · $0.117/M · 42 t/s
    29.2%
    τ²-Bench
  20. 20qwen logo
    qwen3-next-80b-a3b-instruct
    ToolsJSON13.7 intel · $0.090/M · 74 t/s
    21.6%
    τ²-Bench
  21. 21qwen logo
    qwen3-vl-30b-a3b-instruct
    ToolsJSONVision10.0 intel · $0.150/M · 34 t/s
    19.0%
    τ²-Bench

How this is ranked

AI models ranked by Tau²-Bench — multi-turn conversations where the model has to call the right tools, in the right order, against a real API to complete a customer task. This measures whether function calling actually works under pressure, which is a different question from whether a model supports the parameter at all.

Frequently asked

Which Qwen model is best at tool calling?

The best Qwen model for tool calling is Qwen3.6 Plus, completing 97.7% of Tau²-Bench's multi-turn tool-use tasks. Qwen3.6 Max Preview (95.9%) and Qwen3.5 397B A17B (95.6%) round out the top three.

What's a good alternative to Qwen3.6 Plus?

Qwen3.6 Max Preview (95.9%) is the closest alternative on this metric, followed by Qwen3.5 397B A17B (95.6%). See the full ranking above for the tradeoffs.

How many Qwen models are there?

modelgrep tracks 49 Qwen models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Qwen3.7 Max. 21 of them qualify for this ranking.

More Qwen rankings

All rankings