modelgrep

Best Anthropic Models for Math & Science

Match · Updated July 2026

The best Anthropic model for math and science is Claude Fable 5, scoring 92.6% on GPQA Diamond (graduate-level physics, chemistry and biology). Claude Opus 4.8 (92.0%) and Claude Opus 4.7 (91.4%) round out the top three.

92.6%GPQA
59.9Intelligence
$10.00Input /M
1MContext

AI models ranked by GPQA Diamond — graduate-level physics, chemistry and biology questions that can't be answered by lookup. The best LLMs for math, quantitative reasoning and hard-science work.

  1. 1anthropic logo
    claude-fable-5
    ReasoningToolsJSON+159.9 intel · $10.00/M · 1M ctx
    92.6%
    GPQA
  2. 2anthropic logo
    claude-opus-4.8
    ReasoningToolsJSON+155.7 intel · $5.00/M · 1M ctx
    92.0%
    GPQA
  3. 3anthropic logo
    claude-opus-4.7
    ReasoningToolsJSON+153.5 intel · $5.00/M · 1M ctx
    91.4%
    GPQA
  4. 4anthropic logo
    claude-sonnet-5
    ReasoningToolsJSON+153.4 intel · $2.00/M · 1M ctx
    91.1%
    GPQA
  5. 5anthropic logo
    claude-opus-4.6
    ReasoningToolsJSON+137.8 intel · $5.00/M · 1M ctx
    84.0%
    GPQA
  6. 6anthropic logo
    claude-opus-4.5
    ReasoningToolsJSON+134.7 intel · $5.00/M · 200K ctx
    81.0%
    GPQA
  7. 7anthropic logo
    claude-sonnet-4.6
    ReasoningToolsJSON+135.9 intel · $3.00/M · 1M ctx
    79.9%
    GPQA
  8. 8anthropic logo
    claude-sonnet-4.5
    ReasoningToolsJSON+129.3 intel · $3.00/M · 1M ctx
    72.7%
    GPQA
  9. 9anthropic logo
    claude-opus-4
    ReasoningToolsVision25.5 intel · $15.00/M · 200K ctx
    70.1%
    GPQA
  10. 10anthropic logo
    claude-sonnet-4
    ReasoningToolsVision25.5 intel · $3.00/M · 1M ctx
    68.3%
    GPQA
  11. 11anthropic logo
    claude-haiku-4.5
    ReasoningToolsJSON+123.7 intel · $1.00/M · 200K ctx
    64.6%
    GPQA
  12. 12anthropic logo
    claude-3-haiku
    ToolsVision3.9 intel · $0.250/M · 200K ctx
    37.4%
    GPQA

Frequently asked

What is the best Anthropic model for math?

The best Anthropic model for math and science is Claude Fable 5, scoring 92.6% on GPQA Diamond (graduate-level physics, chemistry and biology). Claude Opus 4.8 (92.0%) and Claude Opus 4.7 (91.4%) round out the top three.

What's a good alternative to Claude Fable 5?

Claude Opus 4.8 (92.0%) is the closest alternative on this metric, followed by Claude Opus 4.7 (91.4%). See the full ranking above for the tradeoffs.

How many Anthropic models are there?

modelgrep tracks 15 Anthropic models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Claude Fable 5. 12 of them qualify for this ranking.

More Anthropic rankings

All rankings