The best OpenAI model for reasoning over long inputs is GPT-5.2-Codex, scoring 75.7% on long-context reasoning — a different question from how large a window it accepts. GPT-5 (75.6%) and GPT-5.1 (75.0%) round out the top three.
AI models ranked by LCR — reasoning accuracy over long inputs, not just the size of the window they accept. A large context window is a capacity claim; this is the measurement of whether the model can still reason over information buried deep inside it. Pair it with the longest-context ranking, which sorts on raw window size.
The best OpenAI model for reasoning over long inputs is GPT-5.2-Codex, scoring 75.7% on long-context reasoning — a different question from how large a window it accepts. GPT-5 (75.6%) and GPT-5.1 (75.0%) round out the top three.
GPT-5 (75.6%) is the closest alternative on this metric, followed by GPT-5.1 (75.0%). See the full ranking above for the tradeoffs.
modelgrep tracks 60 OpenAI models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by GPT-5.6 Sol. 25 of them qualify for this ranking.