modelgrep

Best LLMs for Long-Context Reasoning — SpaceXAI

Match · Updated August 2026

The best SpaceXAI model for reasoning over long inputs is Grok 4.5, scoring 67.7% on long-context reasoning — a different question from how large a window it accepts. Grok 4.3 (64.3%) and Grok 4.20 (58.0%) round out the top three.

67.7%LCR
53.8Intelligence
52 t/sSpeed
$2.00Input /M
500KContext
  1. 1x-ai logo
    grok-4.5
    ReasoningToolsJSON+153.8 intel · $2.00/M · 52 t/s
    67.7%
    LCR
  2. 2x-ai logo
    grok-4.3
    ReasoningToolsJSON+137.6 intel · $1.25/M · 95 t/s
    64.3%
    LCR
  3. 3x-ai logo
    grok-4.20
    ReasoningToolsJSON+137.0 intel · $1.25/M · 144 t/s
    58.0%
    LCR

How this is ranked

AI models ranked by LCR — reasoning accuracy over long inputs, not just the size of the window they accept. A large context window is a capacity claim; this is the measurement of whether the model can still reason over information buried deep inside it. Pair it with the longest-context ranking, which sorts on raw window size.

Frequently asked

Which SpaceXAI model reasons best over long inputs?

The best SpaceXAI model for reasoning over long inputs is Grok 4.5, scoring 67.7% on long-context reasoning — a different question from how large a window it accepts. Grok 4.3 (64.3%) and Grok 4.20 (58.0%) round out the top three.

What's a good alternative to Grok 4.5?

Grok 4.3 (64.3%) is the closest alternative on this metric, followed by Grok 4.20 (58.0%). See the full ranking above for the tradeoffs.

How many SpaceXAI models are there?

modelgrep tracks 5 SpaceXAI models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Grok 4.5. 3 of them qualify for this ranking.

More SpaceXAI rankings

All rankings