Nemotron 3 Nano Omni (free) is the best vision-capable NVIDIA model, pairing 14.9 intelligence with image and document understanding. Nemotron 3.5 Content Safety (free) (—) and Nemotron Nano 12B 2 VL (free) (—) round out the top three.
Multimodal large language models that accept image input, ranked by intelligence. The best vision language models (VLMs) for understanding images, documents and charts.
Nemotron 3 Nano Omni (free) is the best vision-capable NVIDIA model, pairing 14.9 intelligence with image and document understanding. Nemotron 3.5 Content Safety (free) (—) and Nemotron Nano 12B 2 VL (free) (—) round out the top three.
Nemotron 3.5 Content Safety (free) (—) is the closest alternative on this metric, followed by Nemotron Nano 12B 2 VL (free) (—). See the full ranking above for the tradeoffs.
modelgrep tracks 10 NVIDIA models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Nemotron 3 Nano Omni (free). 3 of them qualify for this ranking.