modelgrep

Best Meta Vision Models

Match · Updated July 2026

Llama 4 Maverick is the best vision-capable Meta model, pairing 14.3 intelligence with image and document understanding. Llama 4 Scout (10.0) and Llama Guard 4 12B (—) round out the top three.

14.3Intelligence
$0.200Input /M
1.0MContext

Multimodal large language models that accept image input, ranked by intelligence. The best vision language models (VLMs) for understanding images, documents and charts.

  1. 1meta-llama logo
    llama-4-maverick
    ToolsJSONVision14.3 intel · $0.200/M · 1.0M ctx
    14.3
    Intelligence
  2. 2meta-llama logo
    llama-4-scout
    ToolsJSONVision10.0 intel · $0.100/M · 1.3M ctx
    10.0
    Intelligence
  3. 3meta-llama logo
    llama-guard-4-12b
    JSONVision$0.180/M · 1.0M ctx
    Intelligence

Frequently asked

What is the best Meta model for vision?

Llama 4 Maverick is the best vision-capable Meta model, pairing 14.3 intelligence with image and document understanding. Llama 4 Scout (10.0) and Llama Guard 4 12B (—) round out the top three.

What's a good alternative to Llama 4 Maverick?

Llama 4 Scout (10.0) is the closest alternative on this metric, followed by Llama Guard 4 12B (—). See the full ranking above for the tradeoffs.

How many Meta models are there?

modelgrep tracks 8 Meta models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by Llama 4 Maverick. 3 of them qualify for this ranking.

More Meta rankings

All rankings