modelgrep

Best LLMs for Tool Calling — Google

Match · Updated August 2026

The best Google model for tool calling is Gemini 3.1 Pro Preview, completing 95.6% of Tau²-Bench's multi-turn tool-use tasks. Gemini 3.5 Flash (95.3%) and Gemini 2.5 Pro (54.1%) round out the top three.

95.6%τ²-Bench
46.5Intelligence
136 t/sSpeed
$2.00Input /M
1.0MContext
  1. 1google logo
    gemini-3.1-pro-preview
    ReasoningToolsJSON+246.5 intel · $2.00/M · 136 t/s
    95.6%
    τ²-Bench
  2. 2google logo
    gemini-3.5-flash
    ReasoningToolsJSON+250.2 intel · $1.50/M · 275 t/s
    95.3%
    τ²-Bench
  3. 3google logo
    gemini-2.5-pro
    ReasoningToolsJSON+225.8 intel · $1.25/M · 85 t/s
    54.1%
    τ²-Bench
  4. 4google logo
    gemini-3.1-flash-lite
    ReasoningToolsJSON+225.0 intel · $0.250/M · 337 t/s
    31.3%
    τ²-Bench
  5. 5google logo
    gemini-3.1-flash-lite-preview
    ReasoningToolsJSON+225.0 intel · $0.250/M · 337 t/s
    31.3%
    τ²-Bench
  6. 6google logo
    gemini-3.6-flash
    ReasoningToolsJSON+250.1 intel · $1.50/M · 135 t/s
    24.5%
    τ²-Bench
  7. 7google logo
    gemini-2.5-flash-lite
    ReasoningToolsJSON+26.9 intel · $0.100/M · 107 t/s
    19.0%
    τ²-Bench
  8. 8google logo
    gemini-3.5-flash-lite
    ReasoningToolsJSON+236.5 intel · $0.300/M · 160 t/s
    16.5%
    τ²-Bench
  9. 9google logo
    gemini-2.5-flash
    ReasoningToolsJSON+214.1 intel · $0.300/M · 150 t/s
    14.9%
    τ²-Bench

How this is ranked

AI models ranked by Tau²-Bench — multi-turn conversations where the model has to call the right tools, in the right order, against a real API to complete a customer task. This measures whether function calling actually works under pressure, which is a different question from whether a model supports the parameter at all.

Frequently asked

Which Google model is best at tool calling?

The best Google model for tool calling is Gemini 3.1 Pro Preview, completing 95.6% of Tau²-Bench's multi-turn tool-use tasks. Gemini 3.5 Flash (95.3%) and Gemini 2.5 Pro (54.1%) round out the top three.

What's a good alternative to Gemini 3.1 Pro Preview?

Gemini 3.5 Flash (95.3%) is the closest alternative on this metric, followed by Gemini 2.5 Pro (54.1%). See the full ranking above for the tradeoffs.

How many Google models are there?

modelgrep tracks 30 Google models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Gemini 3.5 Flash. 9 of them qualify for this ranking.

More Google rankings

All rankings