The best LLM for instruction following is MiniMax M3, satisfying 82.9% of IFBench's verifiable prompt constraints. Grok 4.3 (81.3%) and Grok 4.20 (81.2%) round out the top three.
AI models ranked by IFBench — how reliably a model obeys explicit, verifiable constraints in the prompt (format, length, inclusion and exclusion rules). High scores mean fewer retries and less prompt-wrangling in production, which often matters more than raw intelligence for pipelines.
The best LLM for instruction following is MiniMax M3, satisfying 82.9% of IFBench's verifiable prompt constraints. Grok 4.3 (81.3%) and Grok 4.20 (81.2%) round out the top three.
The best AI model for instruction following is MiniMax M3, satisfying 82.9% of IFBench's verifiable prompt constraints. Grok 4.3 (81.3%) and Grok 4.20 (81.2%) round out the top three.
Grok 4.3 (81.3%) is the closest alternative on this metric, followed by Grok 4.20 (81.2%). See the full ranking above for the tradeoffs.