modelgrep

Best MiniMax Vision Models

Match · Updated July 2026

MiniMax M3 is the best vision-capable MiniMax model, pairing 44.4 intelligence with image and document understanding. MiniMax M3 (batch) (44.4) and MiniMax-01 (—) round out the top three.

44.4Intelligence
$0.300Input /M
1.0MContext

Multimodal large language models that accept image input, ranked by intelligence. The best vision language models (VLMs) for understanding images, documents and charts.

  1. 1minimax logo
    minimax-m3
    ReasoningToolsJSON+144.4 intel · $0.300/M · 1.0M ctx
    44.4
    Intelligence
  2. 2minimax logo
    minimax-m3:batch
    ReasoningToolsJSON+144.4 intel · $0.150/M · 524K ctx
    44.4
    Intelligence
  3. 3minimax logo
    minimax-01
    Vision$0.200/M · 1.0M ctx
    Intelligence

Frequently asked

What is the best MiniMax model for vision?

MiniMax M3 is the best vision-capable MiniMax model, pairing 44.4 intelligence with image and document understanding. MiniMax M3 (batch) (44.4) and MiniMax-01 (—) round out the top three.

What's a good alternative to MiniMax M3?

MiniMax M3 (batch) (44.4) is the closest alternative on this metric, followed by MiniMax-01 (—). See the full ranking above for the tradeoffs.

How many MiniMax models are there?

modelgrep tracks 9 MiniMax models with live benchmarks, speed, latency and per-provider pricing, led on intelligence by MiniMax M3. 3 of them qualify for this ranking.

More MiniMax rankings

All rankings