all the models — AI benchmark observatory
← Models

DeepSeek-V4.1-Flash

DeepSeek · open weight · deepseek-v4.1-flash

DeepSeek-V4.1-Flash is an MIT-licensed multimodal Mixture-of-Experts model accepting images and text and generating text. It has 552B backbone parameters, 196B Engram conditional-memory parameters, and approximately 763B parameters in the released checkpoint. Its causal encoder-decoder architecture activates 8B parameters per token during prefill and 16B during decode. Trained on 45T multimodal tokens, it supports a 1M-token context, up to 384K output tokens on the DeepSeek API, and continuously adjustable reasoning effort from 1 to 100. CSA2 attention and FP4 KV caching reduce global KV cache storage to 890 bytes per token. The API model name is deepseek-flash. Catalog prices are peak rates per million tokens: $0.30 input, $0.006 cached input, and $1.20 output. Off-peak rates are $0.15, $0.003, and $0.60 respectively, effective September 10, 2026 at 04:00 UTC. Peak hours are 01:00–04:00 and 06:00–10:00 UTC on weekdays; all other times are off-peak (the launch pricing announcement also lists public holidays as off-peak).

CodeForcesGPQATerminal-BenBabyVisionCyberGymDeepSWE 1.1

Benchmark scores

Pricing

  • Fireworks$0.22 / $0.66
  • DeepInfra$0.30 / $1.20
  • Novita$0.30 / $1.20
  • DeepSeek$0.30 / $1.20

Input / output per 1M tokens

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.