all the models — AI benchmark observatory
← Models

DeepSeek-V4-Flash-0423

DeepSeek · open weight · deepseek-v4-flash-0423

DeepSeek-V4-Flash-0423 is the preview release of DeepSeek-V4-Flash, a 284B-parameter MoE model with 13B activated parameters and a 1M-token context window, evaluated here at the default high reasoning effort. It shares the V4 series' hybrid attention architecture combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) for dramatically improved long-context efficiency, Manifold-Constrained Hyper-Connections (mHC) for stable signal propagation, and the Muon optimizer for faster convergence. Pre-trained on more than 32T tokens and post-trained with a two-stage paradigm of domain-specific expert cultivation followed by on-policy distillation, V4-Flash offers reasoning capabilities that closely approach V4-Pro with faster responses and highly cost-effective pricing.

CodeForcesHMMT Feb 26LiveCodeBencGPQAMMLU-ProSWE-Bench Ve

Benchmark scores

Pricing

  • DeepInfra$0.09 / $0.18
  • Novita$0.14 / $0.28

Input / output per 1M tokens

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.
DeepSeek-V4-Flash-0423 Benchmarks · all the models