modelgrep

Best LLMs for Tool Calling

Match · Updated August 2026

The best LLM for tool calling is GLM 5.2, completing 99.1% of Tau²-Bench's multi-turn tool-use tasks. GLM 4.7 Flash (98.8%) and Claude Fable 5 (98.5%) round out the top three.

99.1%τ²-Bench
51.1Intelligence
143 t/sSpeed
$0.760Input /M
1.0MContext
  1. 1z-ai logo
    glm-5.2
    ReasoningToolsJSON51.1 intel · $0.760/M · 143 t/s
    99.1%
    τ²-Bench
  2. 2z-ai logo
    glm-4.7-flash
    ReasoningToolsJSON22.9 intel · $0.060/M · 203K ctx
    98.8%
    τ²-Bench
  3. 3anthropic logo
    claude-fable-5
    ReasoningToolsJSON+159.9 intel · $10.00/M · 64 t/s
    98.5%
    τ²-Bench
  4. 4stepfun logo
    step-3.7-flash
    ReasoningToolsJSON+130.3 intel · $0.200/M · 157 t/s
    98.5%
    τ²-Bench
  5. 5z-ai logo
    glm-5v-turbo
    ReasoningToolsJSON+134.5 intel · $1.20/M · 203K ctx
    98.5%
    τ²-Bench
  6. 6z-ai logo
    glm-5-turbo
    ReasoningToolsJSON38.1 intel · $1.20/M · 203K ctx
    98.5%
    τ²-Bench
  7. 7z-ai logo
    glm-5
    ReasoningToolsJSON39.5 intel · $0.950/M · 205K ctx
    98.2%
    τ²-Bench
  8. 8x-ai logo
    grok-4.3
    ReasoningToolsJSON+137.6 intel · $1.25/M · 1M ctx
    97.7%
    τ²-Bench
  9. 9z-ai logo
    glm-5.1
    ReasoningToolsJSON40.2 intel · $0.966/M · 205K ctx
    97.7%
    τ²-Bench
  10. 10qwen logo
    qwen3.6-plus
    ReasoningToolsJSON+139.6 intel · $0.325/M · 55 t/s
    97.7%
    τ²-Bench
  11. 11deepseek logo
    deepseek-v4-pro
    ReasoningToolsJSON44.3 intel · $0.435/M · 67 t/s
    96.2%
    τ²-Bench
  12. 12qwen logo
    qwen3.6-max-preview
    ReasoningToolsJSON40.0 intel · $1.03/M · 262K ctx
    95.9%
    τ²-Bench
  13. 13moonshotai logo
    kimi-k2.6
    ReasoningToolsJSON+144.2 intel · $0.589/M · 262K ctx
    95.9%
    τ²-Bench
  14. 14moonshotai logo
    kimi-k2.5
    ReasoningToolsJSON+135.4 intel · $0.570/M · 262K ctx
    95.9%
    τ²-Bench
  15. 15z-ai logo
    glm-4.7
    ReasoningToolsJSON33.7 intel · $0.400/M · 541 t/s
    95.9%
    τ²-Bench
  16. 16google logo
    gemini-3.1-pro-preview
    ReasoningToolsJSON+246.5 intel · $2.00/M · 136 t/s
    95.6%
    τ²-Bench
  17. 17qwen logo
    qwen3.5-397b-a17b
    ReasoningToolsJSON+133.7 intel · $0.390/M · 69 t/s
    95.6%
    τ²-Bench
  18. 18google logo
    gemini-3.5-flash
    ReasoningToolsJSON+250.2 intel · $1.50/M · 275 t/s
    95.3%
    τ²-Bench
  19. 19qwen logo
    qwen3.6-35b-a3b
    ReasoningToolsJSON+131.6 intel · $0.140/M · 158 t/s
    95.3%
    τ²-Bench
  20. 20minimax logo
    minimax-m2.5
    ReasoningToolsJSON33.7 intel · $0.150/M · 205K ctx
    95.3%
    τ²-Bench
  21. 21qwen logo
    qwen3.7-max
    ReasoningToolsJSON46.0 intel · $1.48/M · 45 t/s
    94.7%
    τ²-Bench
  22. 22anthropic logo
    claude-opus-4.8
    ReasoningToolsJSON+155.7 intel · $5.00/M · 77 t/s
    94.4%
    τ²-Bench
  23. 23mistralai logo
    mistral-medium-3-5
    ReasoningToolsJSON+129.9 intel · $1.50/M · 152 t/s
    94.2%
    τ²-Bench
  24. 24qwen logo
    qwen3.6-27b
    ReasoningToolsJSON+137.1 intel · $0.289/M · 60 t/s
    94.2%
    τ²-Bench
  25. 25xiaomi logo
    mimo-v2.5-pro
    ReasoningToolsJSON42.2 intel · $0.435/M · 65 t/s
    94.2%
    τ²-Bench

How this is ranked

AI models ranked by Tau²-Bench — multi-turn conversations where the model has to call the right tools, in the right order, against a real API to complete a customer task. This measures whether function calling actually works under pressure, which is a different question from whether a model supports the parameter at all.

Frequently asked

Which LLM is best at tool calling?

The best LLM for tool calling is GLM 5.2, completing 99.1% of Tau²-Bench's multi-turn tool-use tasks. GLM 4.7 Flash (98.8%) and Claude Fable 5 (98.5%) round out the top three.

Which AI model is best at function calling?

The best AI model for tool calling is GLM 5.2, completing 99.1% of Tau²-Bench's multi-turn tool-use tasks. GLM 4.7 Flash (98.8%) and Claude Fable 5 (98.5%) round out the top three.

What's a good alternative to GLM 5.2?

GLM 4.7 Flash (98.8%) is the closest alternative on this metric, followed by Claude Fable 5 (98.5%). See the full ranking above for the tradeoffs.

By maker

All rankings