Qwen3 VL 32B Instruct
Alibaba Cloud / Qwen Team · open weight · qwen3-vl-32b-instruct
Qwen3-VL is a large multimodal model that unifies vision, language, and reasoning to achieve human-level perception and cognition across text, images, and video. Built on a 235B-parameter architecture, it integrates early joint training of visual and textual modalities for strong language grounding. The model supports up to a 1 million-token context window and excels at visual understanding, spatial reasoning, long video comprehension, and tool-based interaction. It can generate code from images, perform precise 2D/3D object grounding, and operate digital interfaces like a visual agent. The “Instruct” version rivals Gemini 2.5 Pro in perception benchmarks, while the “Thinking” version leads in multimodal reasoning and STEM tasks. With multilingual OCR, creative writing, and fine-grained scene interpretation, Qwen3-VL establishes a new open-source frontier for integrated vision-language intelligence.
Benchmark scores
| Benchmark | Score |
|---|---|
| DocVQAtest | 0.97 |
| ScreenSpot | 0.96 |
| CharXiv-D | 0.91 |
| MMLU-Redux | 0.90 |
| AI2D | 0.90 |
| OCRBench | 0.90 |
| InfoVQAtest | 0.87 |
| MMLU | 0.86 |
| IFEval | 0.85 |
| MMLU-Pro | 0.79 |
| BFCL-v3 | 0.70 |
| GPQA | 0.69 |
| AIME 2025 | 0.66 |
| MMMU-Pro | 0.65 |
Pricing
- No provider pricing.
AA metrics
No Artificial Analysis link yet.
Arena Elo
- No Arena snapshot linked.
