all the models — AI benchmark observatory
← Models

Qwen3 VL 30B A3B Instruct

Alibaba Cloud / Qwen Team · open weight · qwen3-vl-30b-a3b-instruct

Qwen3-VL is a large multimodal model that unifies vision, language, and reasoning to achieve human-level perception and cognition across text, images, and video. Built on a 235B-parameter architecture, it integrates early joint training of visual and textual modalities for strong language grounding. The model supports up to a 1 million-token context window and excels at visual understanding, spatial reasoning, long video comprehension, and tool-based interaction. It can generate code from images, perform precise 2D/3D object grounding, and operate digital interfaces like a visual agent. The “Instruct” version rivals Gemini 2.5 Pro in perception benchmarks, while the “Thinking” version leads in multimodal reasoning and STEM tasks. With multilingual OCR, creative writing, and fine-grained scene interpretation, Qwen3-VL establishes a new open-source frontier for integrated vision-language intelligence.

DocVQAtestScreenSpotOCRBenchMMBench-V1.1IFEvalCharXiv-D

Benchmark scores

BenchmarkScore
DocVQAtest0.95
ScreenSpot0.95
OCRBench0.90
MMBench-V1.10.87
IFEval0.86
CharXiv-D0.85
AI2D0.85
MMLU0.85
MMLU-Pro0.78
GPQA0.70
AIME 20250.69
MMMU-Pro0.60

Pricing

  • DeepInfra$0.15 / $0.60

Input / output per 1M tokens

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.
Qwen3 VL 30B A3B Instruct Benchmarks · all the models