modelgrep

Best LLMs for Long-Context Reasoning — Anthropic

Match · Updated August 2026

The best Anthropic model for reasoning over long inputs is Claude Sonnet 5, scoring 70.7% on long-context reasoning — a different question from how large a window it accepts. Claude Opus 4.7 (70.3%) and Claude Opus 5 (70.0%) round out the top three.

70.7%LCR
53.4Intelligence
95 t/sSpeed
$2.00Input /M
1MContext
  1. 1anthropic logo
    claude-sonnet-5
    ReasoningToolsJSON+153.4 intel · $2.00/M · 95 t/s
    70.7%
    LCR
  2. 2anthropic logo
    claude-opus-4.7
    ReasoningToolsJSON+153.5 intel · $5.00/M · 1M ctx
    70.3%
    LCR
  3. 3anthropic logo
    claude-opus-5
    ReasoningToolsJSON+160.7 intel · $5.00/M · 60 t/s
    70.0%
    LCR
  4. 4anthropic logo
    claude-fable-5
    ReasoningToolsJSON+159.9 intel · $10.00/M · 71 t/s
    70.0%
    LCR
  5. 5anthropic logo
    claude-opus-4.8
    ReasoningToolsJSON+155.7 intel · $5.00/M · 1M ctx
    67.7%
    LCR
  6. 6anthropic logo
    claude-opus-4.5
    ReasoningToolsJSON+134.7 intel · $5.00/M · 51 t/s
    65.3%
    LCR
  7. 7anthropic logo
    claude-opus-4.6
    ReasoningToolsJSON+137.8 intel · $5.00/M · 40 t/s
    58.3%
    LCR
  8. 8anthropic logo
    claude-sonnet-4.6
    ReasoningToolsJSON+135.9 intel · $3.00/M · 1M ctx
    57.7%
    LCR
  9. 9anthropic logo
    claude-sonnet-4.5
    ReasoningToolsJSON+129.3 intel · $3.00/M · 1M ctx
    51.3%
    LCR
  10. 10anthropic logo
    claude-sonnet-4
    ReasoningToolsVision25.5 intel · $3.00/M · 1M ctx
    44.3%
    LCR
  11. 11anthropic logo
    claude-haiku-4.5
    ReasoningToolsJSON+123.7 intel · $1.00/M · 91 t/s
    43.7%
    LCR
  12. 12anthropic logo
    claude-opus-4
    ReasoningToolsVision25.5 intel · $15.00/M · 200K ctx
    36.0%
    LCR
  13. 13anthropic logo
    claude-3-haiku
    ToolsVision3.9 intel · $0.250/M · 89 t/s
    21.0%
    LCR

How this is ranked

AI models ranked by LCR — reasoning accuracy over long inputs, not just the size of the window they accept. A large context window is a capacity claim; this is the measurement of whether the model can still reason over information buried deep inside it. Pair it with the longest-context ranking, which sorts on raw window size.

Frequently asked

Which Anthropic model reasons best over long inputs?

The best Anthropic model for reasoning over long inputs is Claude Sonnet 5, scoring 70.7% on long-context reasoning — a different question from how large a window it accepts. Claude Opus 4.7 (70.3%) and Claude Opus 5 (70.0%) round out the top three.

What's a good alternative to Claude Sonnet 5?

Claude Opus 4.7 (70.3%) is the closest alternative on this metric, followed by Claude Opus 5 (70.0%). See the full ranking above for the tradeoffs.

How many Anthropic models are there?

modelgrep tracks 17 Anthropic models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Claude Opus 5. 13 of them qualify for this ranking.

More Anthropic rankings

All rankings