modelgrep

Best LLMs for Long-Context Reasoning — Google

Match · Updated August 2026

The best Google model for reasoning over long inputs is Gemini 3.1 Pro Preview, scoring 72.7% on long-context reasoning — a different question from how large a window it accepts. Gemini 3.6 Flash (69.7%) and Gemini 3.5 Flash (69.3%) round out the top three.

72.7%LCR
46.5Intelligence
136 t/sSpeed
$2.00Input /M
1.0MContext
  1. 1google logo
    gemini-3.1-pro-preview
    ReasoningToolsJSON+246.5 intel · $2.00/M · 136 t/s
    72.7%
    LCR
  2. 2google logo
    gemini-3.6-flash
    ReasoningToolsJSON+250.1 intel · $1.50/M · 135 t/s
    69.7%
    LCR
  3. 3google logo
    gemini-3.5-flash
    ReasoningToolsJSON+250.2 intel · $1.50/M · 270 t/s
    69.3%
    LCR
  4. 4google logo
    gemini-2.5-pro
    ReasoningToolsJSON+225.8 intel · $1.25/M · 85 t/s
    66.0%
    LCR
  5. 5google logo
    gemini-3.1-flash-lite
    ReasoningToolsJSON+225.0 intel · $0.250/M · 135 t/s
    65.3%
    LCR
  6. 6google logo
    gemini-3.1-flash-lite-preview
    ReasoningToolsJSON+225.0 intel · $0.250/M · 337 t/s
    65.3%
    LCR
  7. 7google logo
    gemini-3.5-flash-lite
    ReasoningToolsJSON+236.5 intel · $0.300/M · 160 t/s
    62.0%
    LCR
  8. 8google logo
    gemini-2.5-flash
    ReasoningToolsJSON+214.1 intel · $0.300/M · 150 t/s
    45.9%
    LCR
  9. 9google logo
    gemini-2.5-flash-lite
    ReasoningToolsJSON+26.9 intel · $0.100/M · 107 t/s
    31.3%
    LCR

How this is ranked

AI models ranked by LCR — reasoning accuracy over long inputs, not just the size of the window they accept. A large context window is a capacity claim; this is the measurement of whether the model can still reason over information buried deep inside it. Pair it with the longest-context ranking, which sorts on raw window size.

Frequently asked

Which Google model reasons best over long inputs?

The best Google model for reasoning over long inputs is Gemini 3.1 Pro Preview, scoring 72.7% on long-context reasoning — a different question from how large a window it accepts. Gemini 3.6 Flash (69.7%) and Gemini 3.5 Flash (69.3%) round out the top three.

What's a good alternative to Gemini 3.1 Pro Preview?

Gemini 3.6 Flash (69.7%) is the closest alternative on this metric, followed by Gemini 3.5 Flash (69.3%). See the full ranking above for the tradeoffs.

How many Google models are there?

modelgrep tracks 30 Google models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Gemini 3.5 Flash. 9 of them qualify for this ranking.

More Google rankings

All rankings