← Models
Qwen3.8-Flash-Next
Alibaba Cloud / Qwen Team · open weight · qwen3.8-flash-next
Qwen3.8-Flash-Next is an open-weight experimental preview of the architecture planned for Qwen4, with Hybrid Attention (QSA), Gated Residual, and N-gram Embedding. The card reports 125B parameters with 6B activated, plus 51B n-gram embedding and 4B MTP; Hugging Face BF16 safetensors total 179,999,981,424 parameters (~180B stored). It is a causal LM with a vision encoder for text, image, and video, a native 262,144-token context extensible to 1,000,000 tokens, and thinking on by default (enable_thinking, preserve_thinking, reasoning_effort). This catalog entry is the open-weight checkpoint, not the separate production Qwen3.8-Flash API on Qwen Cloud.
Benchmark scores
| Benchmark | Score |
|---|---|
| MathVision | 0.96 |
| GPQA | 0.92 |
| CharXiv-R | 0.91 |
| RealWorldQA | 0.89 |
| AndroidWorld | 0.84 |
| SWE-bench Multilingual | 0.81 |
| CoWorkBench | 0.74 |
| Toolathlon | 0.73 |
| Humanity's Last Exam | 0.36 |
Pricing
- No provider pricing.
AA metrics
No Artificial Analysis link yet.
Arena Elo
- code1635
