modelgrep

Best LLMs for Long-Context Reasoning

Match · Updated August 2026

The best LLM for reasoning over long inputs is GPT-5.2-Codex, scoring 75.7% on long-context reasoning — a different question from how large a window it accepts. GPT-5 (75.6%) and GPT-5.1 (75.0%) round out the top three.

75.7%LCR
40.1Intelligence
31 t/sSpeed
$1.75Input /M
400KContext
  1. 1openai logo
    gpt-5.2-codex
    ReasoningToolsJSON+140.1 intel · $1.75/M · 31 t/s
    75.7%
    LCR
  2. 2openai logo
    gpt-5
    ReasoningToolsJSON+134.7 intel · $1.25/M · 65 t/s
    75.6%
    LCR
  3. 3openai logo
    gpt-5.1
    ReasoningToolsJSON+136.9 intel · $1.25/M · 60 t/s
    75.0%
    LCR
  4. 4moonshotai logo
    kimi-k3
    ReasoningToolsJSON+157.1 intel · $3.00/M · 64 t/s
    74.7%
    LCR
  5. 5openai logo
    gpt-5.5
    ReasoningToolsJSON+154.8 intel · $5.00/M · 80 t/s
    74.3%
    LCR
  6. 6openai logo
    gpt-5.6-luna
    ReasoningToolsJSON+151.2 intel · $0.100/M · 111 t/s
    74.0%
    LCR
  7. 7openai logo
    gpt-5.6-terra
    ReasoningToolsJSON+155.0 intel · $1.00/M · 83 t/s
    74.0%
    LCR
  8. 8minimax logo
    minimax-m3
    ReasoningToolsJSON+144.4 intel · $0.300/M · 113 t/s
    74.0%
    LCR
  9. 9openai logo
    gpt-5.4
    ReasoningToolsJSON+151.4 intel · $2.50/M · 90 t/s
    74.0%
    LCR
  10. 10openai logo
    gpt-5.3-codex
    ReasoningToolsJSON+144.3 intel · $1.75/M · 61 t/s
    74.0%
    LCR
  11. 11openai logo
    gpt-5.6-sol
    ReasoningToolsJSON+158.9 intel · $5.00/M · 51 t/s
    73.7%
    LCR
  12. 12xiaomi logo
    mimo-v2.5-pro
    ReasoningToolsJSON42.2 intel · $0.435/M · 86 t/s
    73.3%
    LCR
  13. 13google logo
    gemini-3.1-pro-preview
    ReasoningToolsJSON+246.5 intel · $2.00/M · 107 t/s
    72.7%
    LCR
  14. 14openai logo
    gpt-5.2
    ReasoningToolsJSON+142.2 intel · $1.75/M · 81 t/s
    72.7%
    LCR
  15. 15z-ai logo
    glm-5.2
    ReasoningToolsJSON51.1 intel · $0.760/M · 143 t/s
    71.3%
    LCR
  16. 16anthropic logo
    claude-sonnet-5
    ReasoningToolsJSON+153.4 intel · $2.00/M · 79 t/s
    70.7%
    LCR
  17. 17anthropic logo
    claude-opus-4.7
    ReasoningToolsJSON+153.5 intel · $5.00/M · 82 t/s
    70.3%
    LCR
  18. 18anthropic logo
    claude-opus-5
    ReasoningToolsJSON+160.7 intel · $5.00/M · 79 t/s
    70.0%
    LCR
  19. 19anthropic logo
    claude-fable-5
    ReasoningToolsJSON+159.9 intel · $10.00/M · 64 t/s
    70.0%
    LCR
  20. 20google logo
    gemini-3.6-flash
    ReasoningToolsJSON+250.1 intel · $1.50/M · 135 t/s
    69.7%
    LCR
  21. 21qwen logo
    qwen3.6-max-preview
    ReasoningToolsJSON40.0 intel · $1.03/M · 21 t/s
    69.7%
    LCR
  22. 22moonshotai logo
    kimi-k2.6
    ReasoningToolsJSON+144.2 intel · $0.589/M · 141 t/s
    69.7%
    LCR
  23. 23qwen logo
    qwen3.6-plus
    ReasoningToolsJSON+139.6 intel · $0.325/M · 37 t/s
    69.7%
    LCR
  24. 24google logo
    gemini-3.5-flash
    ReasoningToolsJSON+250.2 intel · $1.50/M · 270 t/s
    69.3%
    LCR
  25. 25openai logo
    gpt-5.4-mini
    ReasoningToolsJSON+140.0 intel · $0.750/M · 181 t/s
    69.3%
    LCR

How this is ranked

AI models ranked by LCR — reasoning accuracy over long inputs, not just the size of the window they accept. A large context window is a capacity claim; this is the measurement of whether the model can still reason over information buried deep inside it. Pair it with the longest-context ranking, which sorts on raw window size.

Frequently asked

Which LLM reasons best over long inputs?

The best LLM for reasoning over long inputs is GPT-5.2-Codex, scoring 75.7% on long-context reasoning — a different question from how large a window it accepts. GPT-5 (75.6%) and GPT-5.1 (75.0%) round out the top three.

Which AI model handles long documents best?

The best AI model for reasoning over long inputs is GPT-5.2-Codex, scoring 75.7% on long-context reasoning — a different question from how large a window it accepts. GPT-5 (75.6%) and GPT-5.1 (75.0%) round out the top three.

What's a good alternative to GPT-5.2-Codex?

GPT-5 (75.6%) is the closest alternative on this metric, followed by GPT-5.1 (75.0%). See the full ranking above for the tradeoffs.

By maker

All rankings