The best Anthropic model for instruction following is Claude Fable 5, satisfying 63.5% of IFBench's verifiable prompt constraints. Claude Opus 4.8 (62.2%) and Claude Opus 4.7 (58.6%) round out the top three.
AI models ranked by IFBench — how reliably a model obeys explicit, verifiable constraints in the prompt (format, length, inclusion and exclusion rules). High scores mean fewer retries and less prompt-wrangling in production, which often matters more than raw intelligence for pipelines.
The best Anthropic model for instruction following is Claude Fable 5, satisfying 63.5% of IFBench's verifiable prompt constraints. Claude Opus 4.8 (62.2%) and Claude Opus 4.7 (58.6%) round out the top three.
Claude Opus 4.8 (62.2%) is the closest alternative on this metric, followed by Claude Opus 4.7 (58.6%). See the full ranking above for the tradeoffs.
modelgrep tracks 17 Anthropic models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Claude Opus 5. 11 of them qualify for this ranking.