modelgrep

Best LLMs for Instruction Following — Mistral

Match · Updated August 2026

The best Mistral model for instruction following is Mistral Medium 3.5, satisfying 68.8% of IFBench's verifiable prompt constraints. Mistral Medium 3.1 (39.8%) and Mistral Medium 3 (39.3%) round out the top three.

68.8%IFBench
29.9Intelligence
152 t/sSpeed
$1.50Input /M
262KContext
  1. 1mistralai logo
    mistral-medium-3-5
    ReasoningToolsJSON+129.9 intel · $1.50/M · 152 t/s
    68.8%
    IFBench
  2. 2mistralai logo
    mistral-medium-3.1
    ToolsJSONVision14.7 intel · $0.400/M · 45 t/s
    39.8%
    IFBench
  3. 3mistralai logo
    mistral-medium-3
    ToolsJSONVision12.5 intel · $0.400/M · 41 t/s
    39.3%
    IFBench
  4. 4mistralai logo
    mistral-large-2407
    ToolsJSON7.3 intel · $2.00/M · 131K ctx
    31.6%
    IFBench

How this is ranked

AI models ranked by IFBench — how reliably a model obeys explicit, verifiable constraints in the prompt (format, length, inclusion and exclusion rules). High scores mean fewer retries and less prompt-wrangling in production, which often matters more than raw intelligence for pipelines.

Frequently asked

Which Mistral model follows instructions most reliably?

The best Mistral model for instruction following is Mistral Medium 3.5, satisfying 68.8% of IFBench's verifiable prompt constraints. Mistral Medium 3.1 (39.8%) and Mistral Medium 3 (39.3%) round out the top three.

What's a good alternative to Mistral Medium 3.5?

Mistral Medium 3.1 (39.8%) is the closest alternative on this metric, followed by Mistral Medium 3 (39.3%). See the full ranking above for the tradeoffs.

How many Mistral models are there?

modelgrep tracks 18 Mistral models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Mistral Medium 3.5. 4 of them qualify for this ranking.

More Mistral rankings

All rankings