all the models — AI benchmark observatory
← Models

Qwen3-Next-80B-A3B-Instruct

Alibaba Cloud / Qwen Team · open weight · qwen3-next-80b-a3b-instruct

Qwen3-Next-80B-A3B-Instruct is the first in the Qwen3-Next series, featuring groundbreaking architectural innovations. It uses Hybrid Attention combining Gated DeltaNet and Gated Attention for efficient ultra-long context modeling, High-Sparsity MoE with 512 experts (10 activated + 1 shared) achieving extreme low activation ratio, and Multi-Token Prediction for improved performance and faster inference. With 80B total parameters and only 3B activated, it outperforms Qwen3-32B-Base with 10% training cost and 10x throughput for 32K+ contexts. The model performs on par with Qwen3-235B-A22B-Instruct-2507 while excelling at ultra-long-context tasks up to 256K tokens (extensible to 1M with YaRN). Architecture: 48 layers, 15T training tokens, hybrid layout of 12*(3*(Gated DeltaNet->MoE)->(Gated Attention->MoE)).

MMLU-ReduxIFEvalMMLU-ProGPQABFCL-v3AIME 2025

Benchmark scores

BenchmarkScore
MMLU-Redux0.91
IFEval0.88
MMLU-Pro0.81
GPQA0.73
BFCL-v30.70
AIME 20250.69

Pricing

  • DeepInfra$0.09 / $1.10

Input / output per 1M tokens

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.