The best Anthropic model for math and science is Claude Fable 5, scoring 92.6% on GPQA Diamond (graduate-level physics, chemistry and biology). Claude Opus 4.8 (92.0%) and Claude Opus 4.7 (91.4%) round out the top three.
AI models ranked by GPQA Diamond — graduate-level physics, chemistry and biology questions that can't be answered by lookup. The best LLMs for math, quantitative reasoning and hard-science work.
The best Anthropic model for math and science is Claude Fable 5, scoring 92.6% on GPQA Diamond (graduate-level physics, chemistry and biology). Claude Opus 4.8 (92.0%) and Claude Opus 4.7 (91.4%) round out the top three.
Claude Opus 4.8 (92.0%) is the closest alternative on this metric, followed by Claude Opus 4.7 (91.4%). See the full ranking above for the tradeoffs.
modelgrep tracks 15 Anthropic models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Claude Fable 5. 12 of them qualify for this ranking.