The best OpenAI model for instruction following is GPT-5.2-Codex, satisfying 77.6% of IFBench's verifiable prompt constraints. GPT-5.4 Nano (75.9%) and GPT-5.5 (75.9%) round out the top three.
AI models ranked by IFBench — how reliably a model obeys explicit, verifiable constraints in the prompt (format, length, inclusion and exclusion rules). High scores mean fewer retries and less prompt-wrangling in production, which often matters more than raw intelligence for pipelines.
The best OpenAI model for instruction following is GPT-5.2-Codex, satisfying 77.6% of IFBench's verifiable prompt constraints. GPT-5.4 Nano (75.9%) and GPT-5.5 (75.9%) round out the top three.
GPT-5.4 Nano (75.9%) is the closest alternative on this metric, followed by GPT-5.5 (75.9%). See the full ranking above for the tradeoffs.
modelgrep tracks 60 OpenAI models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by GPT-5.6 Sol. 25 of them qualify for this ranking.