The best Google model for reasoning over long inputs is Gemini 3.1 Pro Preview, scoring 72.7% on long-context reasoning — a different question from how large a window it accepts. Gemini 3.6 Flash (69.7%) and Gemini 3.5 Flash (69.3%) round out the top three.
AI models ranked by LCR — reasoning accuracy over long inputs, not just the size of the window they accept. A large context window is a capacity claim; this is the measurement of whether the model can still reason over information buried deep inside it. Pair it with the longest-context ranking, which sorts on raw window size.
The best Google model for reasoning over long inputs is Gemini 3.1 Pro Preview, scoring 72.7% on long-context reasoning — a different question from how large a window it accepts. Gemini 3.6 Flash (69.7%) and Gemini 3.5 Flash (69.3%) round out the top three.
Gemini 3.6 Flash (69.7%) is the closest alternative on this metric, followed by Gemini 3.5 Flash (69.3%). See the full ranking above for the tradeoffs.
modelgrep tracks 30 Google models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Gemini 3.5 Flash. 9 of them qualify for this ranking.