modelgrep

Best LLMs for Tool Calling — OpenAI

Match · Updated August 2026

The best OpenAI model for tool calling is GPT-5.5, completing 93.9% of Tau²-Bench's multi-turn tool-use tasks. GPT-5.2-Codex (92.1%) and GPT-5.4 (87.1%) round out the top three.

93.9%τ²-Bench
54.8Intelligence
$5.00Input /M
1.1MContext
  1. 1openai logo
    gpt-5.5
    ReasoningToolsJSON+154.8 intel · $5.00/M · 1.1M ctx
    93.9%
    τ²-Bench
  2. 2openai logo
    gpt-5.2-codex
    ReasoningToolsJSON+140.1 intel · $1.75/M · 400K ctx
    92.1%
    τ²-Bench
  3. 3openai logo
    gpt-5.4
    ReasoningToolsJSON+151.4 intel · $2.50/M · 1.1M ctx
    87.1%
    τ²-Bench
  4. 4openai logo
    gpt-5.6-terra
    ReasoningToolsJSON+155.0 intel · $1.00/M · 83 t/s
    86.3%
    τ²-Bench
  5. 5openai logo
    gpt-5.3-codex
    ReasoningToolsJSON+144.3 intel · $1.75/M · 122 t/s
    86.0%
    τ²-Bench
  6. 6openai logo
    gpt-5.6-sol
    ReasoningToolsJSON+158.9 intel · $5.00/M · 51 t/s
    85.1%
    τ²-Bench
  7. 7openai logo
    gpt-5.2
    ReasoningToolsJSON+142.2 intel · $1.75/M · 81 t/s
    84.8%
    τ²-Bench
  8. 8openai logo
    gpt-5
    ReasoningToolsJSON+134.7 intel · $1.25/M · 65 t/s
    84.8%
    τ²-Bench
  9. 9openai logo
    gpt-5.4-mini
    ReasoningToolsJSON+140.0 intel · $0.750/M · 400K ctx
    83.3%
    τ²-Bench
  10. 10openai logo
    gpt-5.1-codex
    ReasoningToolsJSON+134.7 intel · $1.25/M · 65 t/s
    83.0%
    τ²-Bench
  11. 11openai logo
    gpt-5.1
    ReasoningToolsJSON+136.9 intel · $1.25/M · 59 t/s
    81.9%
    τ²-Bench
  12. 12openai logo
    o3
    ReasoningToolsJSON+130.4 intel · $2.00/M · 54 t/s
    80.7%
    τ²-Bench
  13. 13openai logo
    gpt-5.4-nano
    ReasoningToolsJSON+138.2 intel · $0.200/M · 400K ctx
    76.0%
    τ²-Bench
  14. 14openai logo
    gpt-5-mini
    ReasoningToolsJSON+125.3 intel · $0.250/M · 93 t/s
    68.4%
    τ²-Bench
  15. 15openai logo
    gpt-oss-120b
    ReasoningToolsJSON23.8 intel · $0.037/M · 624 t/s
    65.8%
    τ²-Bench
  16. 16openai logo
    gpt-5.1-codex-mini
    ReasoningToolsJSON+130.6 intel · $0.250/M · 61 t/s
    62.9%
    τ²-Bench
  17. 17openai logo
    o1
    ReasoningToolsJSON+123.4 intel · $15.00/M · 200K ctx
    62.6%
    τ²-Bench
  18. 18openai logo
    gpt-oss-20b
    ReasoningToolsJSON14.9 intel · $0.030/M · 236 t/s
    60.2%
    τ²-Bench
  19. 19openai logo
    gpt-oss-20b:free
    ReasoningToolsJSON14.9 intel · Free/M · 17 t/s
    60.2%
    τ²-Bench
  20. 20openai logo
    o4-mini-high
    ReasoningToolsJSON+125.6 intel · $1.10/M · 84 t/s
    55.6%
    τ²-Bench
  21. 21openai logo
    o4-mini
    ReasoningToolsJSON+125.6 intel · $1.10/M · 59 t/s
    55.6%
    τ²-Bench
  22. 22openai logo
    gpt-4.1-mini
    ToolsJSONVision14.8 intel · $0.400/M · 40 t/s
    52.9%
    τ²-Bench
  23. 23openai logo
    gpt-4.1
    ToolsJSONVision19.4 intel · $2.00/M · 64 t/s
    47.1%
    τ²-Bench
  24. 24openai logo
    gpt-5-nano
    ReasoningToolsJSON+119.9 intel · $0.050/M · 98 t/s
    36.5%
    τ²-Bench
  25. 25openai logo
    o3-mini-high
    ReasoningToolsJSON15.6 intel · $1.10/M · 200K ctx
    31.3%
    τ²-Bench

How this is ranked

AI models ranked by Tau²-Bench — multi-turn conversations where the model has to call the right tools, in the right order, against a real API to complete a customer task. This measures whether function calling actually works under pressure, which is a different question from whether a model supports the parameter at all.

Frequently asked

Which OpenAI model is best at tool calling?

The best OpenAI model for tool calling is GPT-5.5, completing 93.9% of Tau²-Bench's multi-turn tool-use tasks. GPT-5.2-Codex (92.1%) and GPT-5.4 (87.1%) round out the top three.

What's a good alternative to GPT-5.5?

GPT-5.2-Codex (92.1%) is the closest alternative on this metric, followed by GPT-5.4 (87.1%). See the full ranking above for the tradeoffs.

How many OpenAI models are there?

modelgrep tracks 60 OpenAI models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by GPT-5.6 Sol. 25 of them qualify for this ranking.

More OpenAI rankings

All rankings