modelgrep

Best LLMs for Tool Calling — Z.ai

Match · Updated August 2026

The best Z.ai model for tool calling is GLM 5.2, completing 99.1% of Tau²-Bench's multi-turn tool-use tasks. GLM 4.7 Flash (98.8%) and GLM 5V Turbo (98.5%) round out the top three.

99.1%τ²-Bench
51.1Intelligence
156 t/sSpeed
$0.760Input /M
1.0MContext
  1. 1z-ai logo
    glm-5.2
    ReasoningToolsJSON51.1 intel · $0.760/M · 156 t/s
    99.1%
    τ²-Bench
  2. 2z-ai logo
    glm-4.7-flash
    ReasoningToolsJSON22.9 intel · $0.060/M · 43 t/s
    98.8%
    τ²-Bench
  3. 3z-ai logo
    glm-5v-turbo
    ReasoningToolsJSON+134.5 intel · $1.20/M · 203K ctx
    98.5%
    τ²-Bench
  4. 4z-ai logo
    glm-5-turbo
    ReasoningToolsJSON38.1 intel · $1.20/M · 203K ctx
    98.5%
    τ²-Bench
  5. 5z-ai logo
    glm-5
    ReasoningToolsJSON39.5 intel · $0.950/M · 205K ctx
    98.2%
    τ²-Bench
  6. 6z-ai logo
    glm-5.1
    ReasoningToolsJSON40.2 intel · $0.966/M · 205K ctx
    97.7%
    τ²-Bench
  7. 7z-ai logo
    glm-4.7
    ReasoningToolsJSON33.7 intel · $0.400/M · 496 t/s
    95.9%
    τ²-Bench
  8. 8z-ai logo
    glm-4.6
    ReasoningToolsJSON23.0 intel · $0.500/M · 39 t/s
    76.9%
    τ²-Bench
  9. 9z-ai logo
    glm-4.5-air
    ReasoningTools16.5 intel · $0.130/M · 42 t/s
    46.5%
    τ²-Bench
  10. 10z-ai logo
    glm-4.5
    ReasoningToolsJSON19.5 intel · $0.600/M · 33 t/s
    43.0%
    τ²-Bench
  11. 11z-ai logo
    glm-4.6v
    ReasoningToolsJSON+111.0 intel · $0.300/M · 33 t/s
    30.7%
    τ²-Bench
  12. 12z-ai logo
    glm-4.5v
    ReasoningToolsJSON+17.0 intel · $0.600/M · 39 t/s
    19.6%
    τ²-Bench

How this is ranked

AI models ranked by Tau²-Bench — multi-turn conversations where the model has to call the right tools, in the right order, against a real API to complete a customer task. This measures whether function calling actually works under pressure, which is a different question from whether a model supports the parameter at all.

Frequently asked

Which Z.ai model is best at tool calling?

The best Z.ai model for tool calling is GLM 5.2, completing 99.1% of Tau²-Bench's multi-turn tool-use tasks. GLM 4.7 Flash (98.8%) and GLM 5V Turbo (98.5%) round out the top three.

What's a good alternative to GLM 5.2?

GLM 4.7 Flash (98.8%) is the closest alternative on this metric, followed by GLM 5V Turbo (98.5%). See the full ranking above for the tradeoffs.

How many Z.ai models are there?

modelgrep tracks 12 Z.ai models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by GLM 5.2. 12 of them qualify for this ranking.

More Z.ai rankings

All rankings