← Models
Qwen2.5 VL 72B Instruct
Alibaba Cloud / Qwen Team · open weight · qwen2.5-vl-72b
Qwen2.5-VL is the new flagship vision-language model of Qwen, significantly improved from Qwen2-VL. It excels at recognizing objects, analyzing text/charts/layouts in images, acting as a visual agent, understanding long videos (over 1 hour) with event pinpointing, performing visual localization (bounding boxes/points), and generating structured outputs from documents.
Benchmark scores
| Benchmark | Score |
|---|---|
| DocVQA | 0.96 |
| Android Control Low_EM | 0.94 |
| ChartQA | 0.90 |
| OCRBench | 0.89 |
| AI2D | 0.88 |
| MMBench | 0.88 |
| ScreenSpot | 0.87 |
| AITZ_EM | 0.83 |
| MMMU | 0.70 |
| MMMU-Pro | 0.51 |
Pricing
- No provider pricing.
AA metrics
No Artificial Analysis link yet.
Arena Elo
- No Arena snapshot linked.
