all the models — AI benchmark observatory
← Models

Qwen3.8-Flash-Next

Alibaba Cloud / Qwen Team · open weight · qwen3.8-flash-next

Qwen3.8-Flash-Next is an open-weight experimental preview of the architecture planned for Qwen4, with Hybrid Attention (QSA), Gated Residual, and N-gram Embedding. The card reports 125B parameters with 6B activated, plus 51B n-gram embedding and 4B MTP; Hugging Face BF16 safetensors total 179,999,981,424 parameters (~180B stored). It is a causal LM with a vision encoder for text, image, and video, a native 262,144-token context extensible to 1,000,000 tokens, and thinking on by default (enable_thinking, preserve_thinking, reasoning_effort). This catalog entry is the open-weight checkpoint, not the separate production Qwen3.8-Flash API on Qwen Cloud.

MathVisionGPQACharXiv-RRealWorldQAAndroidWorldSWE-bench Mu

Benchmark scores

Pricing

  • No provider pricing.

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • code1635
Qwen3.8-Flash-Next Benchmarks · all the models