modelgrep

Best LLMs for Tool Calling — Anthropic

Match · Updated August 2026

The best Anthropic model for tool calling is Claude Fable 5, completing 98.5% of Tau²-Bench's multi-turn tool-use tasks. Claude Opus 4.8 (94.4%) and Claude Opus 4.7 (88.6%) round out the top three.

98.5%τ²-Bench
59.9Intelligence
82 t/sSpeed
$10.00Input /M
1MContext
  1. 1anthropic logo
    claude-fable-5
    ReasoningToolsJSON+159.9 intel · $10.00/M · 82 t/s
    98.5%
    τ²-Bench
  2. 2anthropic logo
    claude-opus-4.8
    ReasoningToolsJSON+155.7 intel · $5.00/M · 1M ctx
    94.4%
    τ²-Bench
  3. 3anthropic logo
    claude-opus-4.7
    ReasoningToolsJSON+153.5 intel · $5.00/M · 1M ctx
    88.6%
    τ²-Bench
  4. 4anthropic logo
    claude-opus-4.5
    ReasoningToolsJSON+134.7 intel · $5.00/M · 52 t/s
    86.3%
    τ²-Bench
  5. 5anthropic logo
    claude-opus-4.6
    ReasoningToolsJSON+137.8 intel · $5.00/M · 1M ctx
    84.8%
    τ²-Bench
  6. 6anthropic logo
    claude-sonnet-4.6
    ReasoningToolsJSON+135.9 intel · $3.00/M · 1M ctx
    79.5%
    τ²-Bench
  7. 7anthropic logo
    claude-sonnet-4.5
    ReasoningToolsJSON+129.3 intel · $3.00/M · 36 t/s
    70.5%
    τ²-Bench
  8. 8anthropic logo
    claude-sonnet-4
    ReasoningToolsVision25.5 intel · $3.00/M · 46 t/s
    52.3%
    τ²-Bench
  9. 9anthropic logo
    claude-haiku-4.5
    ReasoningToolsJSON+123.7 intel · $1.00/M · 91 t/s
    32.5%
    τ²-Bench
  10. 10anthropic logo
    claude-opus-5
    ReasoningToolsJSON+160.7 intel · $5.00/M · 60 t/s
    30.3%
    τ²-Bench
  11. 11anthropic logo
    claude-sonnet-5
    ReasoningToolsJSON+153.4 intel · $2.00/M · 95 t/s
    28.2%
    τ²-Bench
  12. 12anthropic logo
    claude-3-haiku
    ToolsVision3.9 intel · $0.250/M · 200K ctx
    21.1%
    τ²-Bench

How this is ranked

AI models ranked by Tau²-Bench — multi-turn conversations where the model has to call the right tools, in the right order, against a real API to complete a customer task. This measures whether function calling actually works under pressure, which is a different question from whether a model supports the parameter at all.

Frequently asked

Which Anthropic model is best at tool calling?

The best Anthropic model for tool calling is Claude Fable 5, completing 98.5% of Tau²-Bench's multi-turn tool-use tasks. Claude Opus 4.8 (94.4%) and Claude Opus 4.7 (88.6%) round out the top three.

What's a good alternative to Claude Fable 5?

Claude Opus 4.8 (94.4%) is the closest alternative on this metric, followed by Claude Opus 4.7 (88.6%). See the full ranking above for the tradeoffs.

How many Anthropic models are there?

modelgrep tracks 17 Anthropic models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Claude Opus 5. 12 of them qualify for this ranking.

More Anthropic rankings

All rankings