GLM 5V Turbo is the best vision-capable Z.ai model, pairing 34.5 intelligence with image and document understanding. GLM 4.6V (11.0) and GLM 4.5V (7.0) round out the top three.
Multimodal large language models that accept image input, ranked by intelligence. The best vision language models (VLMs) for understanding images, documents and charts.
GLM 5V Turbo is the best vision-capable Z.ai model, pairing 34.5 intelligence with image and document understanding. GLM 4.6V (11.0) and GLM 4.5V (7.0) round out the top three.
GLM 4.6V (11.0) is the closest alternative on this metric, followed by GLM 4.5V (7.0). See the full ranking above for the tradeoffs.
modelgrep tracks 12 Z.ai models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by GLM 5.2. 3 of them qualify for this ranking.