qwen/qwen3-vl-32b-instruct
Qwen3-VL-32B-Instruct is a large-scale multimodal vision-language model designed for high-precision understanding and reasoning across text, images, and video. With 32 billion parameters, it combines deep visual perception with advanced text...
| Provider | In $/M | Out $/M | Context | Uptime |
|---|---|---|---|---|
| Alibabafp8 | $0.104 | $0.416 | 131K | 100% |
Qwen3 VL 32B Instruct costs $0.104 per million input tokens and $0.416 per million output tokens via OpenRouter, making it 60th cheapest of 332 paid models.
Qwen3 VL 32B Instruct scores 11.1 on the Artificial Analysis Intelligence Index, ranking 150th of 182 benchmarked models, with a GPQA Diamond score of 67%.
Qwen3 VL 32B Instruct supports a 131K-token context window and can output up to 33K tokens. It accepts text, image input.