Qwen3.6 Plus is the best vision-capable Qwen model, pairing 39.6 intelligence with image and document understanding. Qwen3.7 Plus (39.0) and Qwen3.6 27B (37.1) round out the top three.
Multimodal large language models that accept image input, ranked by intelligence. The best vision language models (VLMs) for understanding images, documents and charts.
Qwen3.6 Plus is the best vision-capable Qwen model, pairing 39.6 intelligence with image and document understanding. Qwen3.7 Plus (39.0) and Qwen3.6 27B (37.1) round out the top three.
Qwen3.7 Plus (39.0) is the closest alternative on this metric, followed by Qwen3.6 27B (37.1). See the full ranking above for the tradeoffs.
modelgrep tracks 48 Qwen models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Qwen3.7 Max. 22 of them qualify for this ranking.