DeepSeek-V4.1-Flash
DeepSeek · open weight · deepseek-v4.1-flash
DeepSeek-V4.1-Flash is an MIT-licensed multimodal Mixture-of-Experts model accepting images and text and generating text. It has 552B backbone parameters, 196B Engram conditional-memory parameters, and approximately 763B parameters in the released checkpoint. Its causal encoder-decoder architecture activates 8B parameters per token during prefill and 16B during decode. Trained on 45T multimodal tokens, it supports a 1M-token context, up to 384K output tokens on the DeepSeek API, and continuously adjustable reasoning effort from 1 to 100. CSA2 attention and FP4 KV caching reduce global KV cache storage to 890 bytes per token. The API model name is deepseek-flash. Catalog prices are peak rates per million tokens: $0.30 input, $0.006 cached input, and $1.20 output. Off-peak rates are $0.15, $0.003, and $0.60 respectively, effective September 10, 2026 at 04:00 UTC. Peak hours are 01:00–04:00 and 06:00–10:00 UTC on weekdays; all other times are off-peak (the launch pricing announcement also lists public holidays as off-peak).
Benchmark scores
| Benchmark | Score |
|---|---|
| CodeForces | 3471.00 |
| GPQA | 0.91 |
| Terminal-Bench 2.1 | 0.91 |
| BabyVision | 0.90 |
| CyberGym | 0.88 |
| DeepSWE 1.1 | 0.74 |
| Humanity's Last Exam | 0.37 |
Pricing
- Fireworks$0.22 / $0.66
- DeepInfra$0.30 / $1.20
- Novita$0.30 / $1.20
- DeepSeek$0.30 / $1.20
Input / output per 1M tokens
AA metrics
No Artificial Analysis link yet.
Arena Elo
- No Arena snapshot linked.
