The best Google model for instruction following is Gemini 3.1 Flash Lite, satisfying 77.2% of IFBench's verifiable prompt constraints. Gemini 3.1 Flash Lite Preview (77.2%) and Gemini 3.1 Pro Preview (77.1%) round out the top three.
AI models ranked by IFBench — how reliably a model obeys explicit, verifiable constraints in the prompt (format, length, inclusion and exclusion rules). High scores mean fewer retries and less prompt-wrangling in production, which often matters more than raw intelligence for pipelines.
The best Google model for instruction following is Gemini 3.1 Flash Lite, satisfying 77.2% of IFBench's verifiable prompt constraints. Gemini 3.1 Flash Lite Preview (77.2%) and Gemini 3.1 Pro Preview (77.1%) round out the top three.
Gemini 3.1 Flash Lite Preview (77.2%) is the closest alternative on this metric, followed by Gemini 3.1 Pro Preview (77.1%). See the full ranking above for the tradeoffs.
modelgrep tracks 30 Google models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Gemini 3.5 Flash. 7 of them qualify for this ranking.