modelgrep

Longest-Context NVIDIA Models

Match · Updated July 2026

Nemotron 3 Ultra (free) has the largest context window of any NVIDIA model, at 1M tokens. Nemotron 3 Super (1M) and Nemotron 3 Ultra (512K) round out the top three.

1MContext
FreeInput /M

AI models with the largest context windows, ranked by token capacity. The best large language models for long documents, codebases and extended conversations.

  1. 1nvidia logo
    nemotron-3-ultra-550b-a55b:free
    ReasoningToolsFree/M
    1M
    Context
  2. 2nvidia logo
    nemotron-3-super-120b-a12b
    ReasoningToolsJSON$0.085/M
    1M
    Context
  3. 3nvidia logo
    nemotron-3-ultra-550b-a55b
    ReasoningToolsJSON$0.600/M
    512K
    Context
  4. 4nvidia logo
    nemotron-3-super-120b-a12b:free
    ReasoningToolsJSONFree/M
    262K
    Context
  5. 5nvidia logo
    nemotron-3-nano-30b-a3b
    ReasoningToolsJSON7.4 intel · $0.050/M
    262K
    Context
  6. 6nvidia logo
    nemotron-3-nano-omni-30b-a3b-reasoning:free
    ReasoningToolsVision+114.9 intel · Free/M
    256K
    Context
  7. 7nvidia logo
    nemotron-3-nano-30b-a3b:free
    ReasoningToolsFree/M
    256K
    Context

Frequently asked

Which NVIDIA model has the largest context window?

Nemotron 3 Ultra (free) has the largest context window of any NVIDIA model, at 1M tokens. Nemotron 3 Super (1M) and Nemotron 3 Ultra (512K) round out the top three.

What's a good alternative to Nemotron 3 Ultra (free)?

Nemotron 3 Super (1M) is the closest alternative on this metric, followed by Nemotron 3 Ultra (512K). See the full ranking above for the tradeoffs.

How many NVIDIA models are there?

modelgrep tracks 10 NVIDIA models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Nemotron 3 Nano Omni (free). 7 of them qualify for this ranking.

More NVIDIA rankings

All rankings