The best Meta model for instruction following is Llama 4 Maverick, satisfying 43.0% of IFBench's verifiable prompt constraints. Llama 4 Scout (39.5%) is next.
AI models ranked by IFBench — how reliably a model obeys explicit, verifiable constraints in the prompt (format, length, inclusion and exclusion rules). High scores mean fewer retries and less prompt-wrangling in production, which often matters more than raw intelligence for pipelines.
The best Meta model for instruction following is Llama 4 Maverick, satisfying 43.0% of IFBench's verifiable prompt constraints. Llama 4 Scout (39.5%) is next.
Llama 4 Scout (39.5%) is the closest alternative on this metric. See the full ranking above for the tradeoffs.
modelgrep tracks 8 Meta models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Llama 4 Maverick. 2 of them qualify for this ranking.