GPT-5.6 Sol is the best vision-capable OpenAI model, pairing 58.9 intelligence with image and document understanding. GPT-5.6 Terra (55.0) and GPT-5.5 (54.8) round out the top three.
Multimodal large language models that accept image input, ranked by intelligence. The best vision language models (VLMs) for understanding images, documents and charts.
GPT-5.6 Sol is the best vision-capable OpenAI model, pairing 58.9 intelligence with image and document understanding. GPT-5.6 Terra (55.0) and GPT-5.5 (54.8) round out the top three.
GPT-5.6 Terra (55.0) is the closest alternative on this metric, followed by GPT-5.5 (54.8). See the full ranking above for the tradeoffs.
modelgrep tracks 70 OpenAI models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by GPT-5.6 Sol. 25 of them qualify for this ranking.