modelgrep

Best LLMs for Tool Calling — Mistral

Match · Updated August 2026

The best Mistral model for tool calling is Mistral Medium 3.5, completing 94.2% of Tau²-Bench's multi-turn tool-use tasks. Mistral Medium 3.1 (40.6%) and Mistral Large 2407 (33.0%) round out the top three.

94.2%τ²-Bench
29.9Intelligence
152 t/sSpeed
$1.50Input /M
262KContext
  1. 1mistralai logo
    mistral-medium-3-5
    ReasoningToolsJSON+129.9 intel · $1.50/M · 152 t/s
    94.2%
    τ²-Bench
  2. 2mistralai logo
    mistral-medium-3.1
    ToolsJSONVision14.7 intel · $0.400/M · 45 t/s
    40.6%
    τ²-Bench
  3. 3mistralai logo
    mistral-large-2407
    ToolsJSON7.3 intel · $2.00/M · 131K ctx
    33.0%
    τ²-Bench
  4. 4mistralai logo
    mistral-medium-3
    ToolsJSONVision12.5 intel · $0.400/M · 41 t/s
    24.3%
    τ²-Bench

How this is ranked

AI models ranked by Tau²-Bench — multi-turn conversations where the model has to call the right tools, in the right order, against a real API to complete a customer task. This measures whether function calling actually works under pressure, which is a different question from whether a model supports the parameter at all.

Frequently asked

Which Mistral model is best at tool calling?

The best Mistral model for tool calling is Mistral Medium 3.5, completing 94.2% of Tau²-Bench's multi-turn tool-use tasks. Mistral Medium 3.1 (40.6%) and Mistral Large 2407 (33.0%) round out the top three.

What's a good alternative to Mistral Medium 3.5?

Mistral Medium 3.1 (40.6%) is the closest alternative on this metric, followed by Mistral Large 2407 (33.0%). See the full ranking above for the tradeoffs.

How many Mistral models are there?

modelgrep tracks 18 Mistral models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Mistral Medium 3.5. 4 of them qualify for this ranking.

More Mistral rankings

All rankings