all the models — AI benchmark observatory
← Models

Qwen3 VL 235B A22B Thinking

Alibaba Cloud / Qwen Team · open weight · qwen3-vl-235b-a22b-thinking

Qwen3-VL-235B-A22B-Thinking is the most powerful vision-language model in the Qwen series, featuring 236B parameters with MoE architecture for reasoning-enhanced multimodal understanding. Key capabilities include: Visual Agent (operates PC/mobile GUIs, recognizes elements, invokes tools), Visual Coding (generates Draw.io/HTML/CSS/JS from images/videos), Advanced Spatial Perception (2D grounding and 3D grounding for spatial reasoning and embodied AI), Long Context & Video Understanding (native 256K context expandable to 1M, handles hours-long video with second-level indexing), Enhanced Multimodal Reasoning (excels in STEM/Math with causal analysis), Upgraded Visual Recognition (celebrities, anime, products, landmarks, flora/fauna), and Expanded OCR (32 languages, robust in low light/blur/tilt). Architecture innovations include Interleaved-MRoPE for positional embeddings, DeepStack for multi-level ViT feature fusion, and Text-Timestamp Alignment for precise video temporal modeling.

ZebraLogicDocVQAtestScreenSpotCountBenchMMLU-ReduxDesign2Code

Benchmark scores

Pricing

  • No provider pricing.

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.
Qwen3 VL 235B A22B Thinking Benchmarks · all the models