modelgrep

Best LLMs for Instruction Following — Anthropic

Match · Updated August 2026

The best Anthropic model for instruction following is Claude Fable 5, satisfying 63.5% of IFBench's verifiable prompt constraints. Claude Opus 4.8 (62.2%) and Claude Opus 4.7 (58.6%) round out the top three.

63.5%IFBench
59.9Intelligence
82 t/sSpeed
$10.00Input /M
1MContext
  1. 1anthropic logo
    claude-fable-5
    ReasoningToolsJSON+159.9 intel · $10.00/M · 82 t/s
    63.5%
    IFBench
  2. 2anthropic logo
    claude-opus-4.8
    ReasoningToolsJSON+155.7 intel · $5.00/M · 1M ctx
    62.2%
    IFBench
  3. 3anthropic logo
    claude-opus-4.7
    ReasoningToolsJSON+153.5 intel · $5.00/M · 1M ctx
    58.6%
    IFBench
  4. 4anthropic logo
    claude-sonnet-4
    ReasoningToolsVision25.5 intel · $3.00/M · 46 t/s
    45.4%
    IFBench
  5. 5anthropic logo
    claude-opus-4.6
    ReasoningToolsJSON+137.8 intel · $5.00/M · 1M ctx
    44.6%
    IFBench
  6. 6anthropic logo
    claude-opus-4
    ReasoningToolsVision25.5 intel · $15.00/M · 23 t/s
    43.3%
    IFBench
  7. 7anthropic logo
    claude-opus-4.5
    ReasoningToolsJSON+134.7 intel · $5.00/M · 52 t/s
    43.0%
    IFBench
  8. 8anthropic logo
    claude-sonnet-4.5
    ReasoningToolsJSON+129.3 intel · $3.00/M · 36 t/s
    42.7%
    IFBench
  9. 9anthropic logo
    claude-haiku-4.5
    ReasoningToolsJSON+123.7 intel · $1.00/M · 91 t/s
    42.0%
    IFBench
  10. 10anthropic logo
    claude-sonnet-4.6
    ReasoningToolsJSON+135.9 intel · $3.00/M · 1M ctx
    41.2%
    IFBench
  11. 11anthropic logo
    claude-3-haiku
    ToolsVision3.9 intel · $0.250/M · 200K ctx
    36.1%
    IFBench

How this is ranked

AI models ranked by IFBench — how reliably a model obeys explicit, verifiable constraints in the prompt (format, length, inclusion and exclusion rules). High scores mean fewer retries and less prompt-wrangling in production, which often matters more than raw intelligence for pipelines.

Frequently asked

Which Anthropic model follows instructions most reliably?

The best Anthropic model for instruction following is Claude Fable 5, satisfying 63.5% of IFBench's verifiable prompt constraints. Claude Opus 4.8 (62.2%) and Claude Opus 4.7 (58.6%) round out the top three.

What's a good alternative to Claude Fable 5?

Claude Opus 4.8 (62.2%) is the closest alternative on this metric, followed by Claude Opus 4.7 (58.6%). See the full ranking above for the tradeoffs.

How many Anthropic models are there?

modelgrep tracks 17 Anthropic models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Claude Opus 5. 11 of them qualify for this ranking.

More Anthropic rankings

All rankings