The best Mistral model for instruction following is Mistral Medium 3.5, satisfying 68.8% of IFBench's verifiable prompt constraints. Mistral Medium 3.1 (39.8%) and Mistral Medium 3 (39.3%) round out the top three.
AI models ranked by IFBench — how reliably a model obeys explicit, verifiable constraints in the prompt (format, length, inclusion and exclusion rules). High scores mean fewer retries and less prompt-wrangling in production, which often matters more than raw intelligence for pipelines.
The best Mistral model for instruction following is Mistral Medium 3.5, satisfying 68.8% of IFBench's verifiable prompt constraints. Mistral Medium 3.1 (39.8%) and Mistral Medium 3 (39.3%) round out the top three.
Mistral Medium 3.1 (39.8%) is the closest alternative on this metric, followed by Mistral Medium 3 (39.3%). See the full ranking above for the tradeoffs.
modelgrep tracks 18 Mistral models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Mistral Medium 3.5. 4 of them qualify for this ranking.