DeepSeek-V4-Pro-Max
DeepSeek · open weight · deepseek-v4-pro-max
DeepSeek-V4-Pro-Max is the maximum reasoning effort mode of DeepSeek-V4-Pro, a 1.6T-parameter MoE model with 49B activated parameters and a 1M-token context window. It introduces a hybrid attention architecture combining Compressed Sparse Attention (CSA) and Heavily Compressed Attention (HCA) for dramatically improved long-context efficiency, requiring only 27% of single-token inference FLOPs and 10% of KV cache compared with DeepSeek-V3.2 at 1M-token context. The model also incorporates Manifold-Constrained Hyper-Connections (mHC) for stable signal propagation and is trained with the Muon optimizer for faster convergence. Pre-trained on more than 32T tokens, V4-Pro-Max significantly advances open-source knowledge capabilities, achieves top-tier performance in coding benchmarks, and bridges the gap with leading closed-source models on reasoning and agentic tasks.
Benchmark scores
| Benchmark | Score |
|---|---|
| CodeForces | 1.00 |
| HMMT Feb 26 | 0.95 |
| LiveCodeBench | 0.94 |
| MathArena Apex | 0.90 |
| GPQA | 0.90 |
| IMO-AnswerBench | 0.90 |
| MMLU-Pro | 0.88 |
| BrowseComp | 0.83 |
| SWE-Bench Verified | 0.81 |
| SWE-bench Multilingual | 0.76 |
| MCP Atlas | 0.74 |
| Humanity's Last Exam | 0.48 |
Pricing
- DeepInfra$1.30 / $2.60
Input / output per 1M tokens
AA metrics
No Artificial Analysis link yet.
Arena Elo
- No Arena snapshot linked.
