← Models
Qwen2-VL-72B-Instruct
Alibaba Cloud / Qwen Team · open weight · qwen2-vl-72b
An instruction-tuned, large multimodal model that excels at visual understanding and step-by-step reasoning. It supports image and video input, with dynamic resolution handling and improved positional embeddings (M-ROPE), enabling advanced capabilities such as complex problem solving, multilingual text recognition in images, and agent-like interactions in video contexts.
Benchmark scores
| Benchmark | Score |
|---|---|
| DocVQAtest | 0.96 |
| VCR_en_easy | 0.92 |
| ChartQA | 0.88 |
| OCRBench | 0.88 |
| MMBench | 0.86 |
| TextVQA | 0.85 |
| MMMU-Pro | 0.46 |
Pricing
- No provider pricing.
AA metrics
No Artificial Analysis link yet.
Arena Elo
- No Arena snapshot linked.
