Qwen3 VL 235B A22B Thinking
Alibaba Cloud / Qwen Team · open weight · qwen3-vl-235b-a22b-thinking
Qwen3-VL-235B-A22B-Thinking is the most powerful vision-language model in the Qwen series, featuring 236B parameters with MoE architecture for reasoning-enhanced multimodal understanding. Key capabilities include: Visual Agent (operates PC/mobile GUIs, recognizes elements, invokes tools), Visual Coding (generates Draw.io/HTML/CSS/JS from images/videos), Advanced Spatial Perception (2D grounding and 3D grounding for spatial reasoning and embodied AI), Long Context & Video Understanding (native 256K context expandable to 1M, handles hours-long video with second-level indexing), Enhanced Multimodal Reasoning (excels in STEM/Math with causal analysis), Upgraded Visual Recognition (celebrities, anime, products, landmarks, flora/fauna), and Expanded OCR (32 languages, robust in low light/blur/tilt). Architecture innovations include Interleaved-MRoPE for positional embeddings, DeepStack for multi-level ViT feature fusion, and Text-Timestamp Alignment for precise video temporal modeling.
Benchmark scores
| Benchmark | Score |
|---|---|
| ZebraLogic | 0.97 |
| DocVQAtest | 0.96 |
| ScreenSpot | 0.95 |
| CountBench | 0.94 |
| MMLU-Redux | 0.94 |
| Design2Code | 0.93 |
| MIABench | 0.93 |
| RefCOCO-avg | 0.92 |
| MMBench-V1.1 | 0.91 |
| MMLU | 0.91 |
| AIME 2025 | 0.90 |
| InfoVQAtest | 0.90 |
| AI2D | 0.89 |
| IFEval | 0.88 |
| OCRBench | 0.88 |
| MathVista-Mini | 0.86 |
| MathVerse-Mini | 0.85 |
| MMLU-Pro | 0.84 |
| SIFO | 0.77 |
| BFCL-v3 | 0.72 |
| SIFO-Multiturn | 0.71 |
| MMMU-Pro | 0.69 |
| Humanity's Last Exam | 0.14 |
Pricing
- No provider pricing.
AA metrics
No Artificial Analysis link yet.
Arena Elo
- No Arena snapshot linked.
