Kimi K3 is the best vision-capable MoonshotAI model, pairing 57.1 intelligence with image and document understanding. Kimi K2.6 (44.2) and Kimi K2.7 Code (41.9) round out the top three.
Multimodal large language models that accept image input, ranked by intelligence. The best vision language models (VLMs) for understanding images, documents and charts.
Kimi K3 is the best vision-capable MoonshotAI model, pairing 57.1 intelligence with image and document understanding. Kimi K2.6 (44.2) and Kimi K2.7 Code (41.9) round out the top three.
Kimi K2.6 (44.2) is the closest alternative on this metric, followed by Kimi K2.7 Code (41.9). See the full ranking above for the tradeoffs.
modelgrep tracks 7 MoonshotAI models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Kimi K3. 4 of them qualify for this ranking.