DeepSeek-V3.1
DeepSeek · open weight · deepseek-v3.1
DeepSeek-V3.1 is a hybrid model supporting both thinking and non-thinking modes through different chat templates. Built on DeepSeek-V3.1-Base with a two-phase long context extension (32K phase: 630B tokens, 128K phase: 209B tokens), it features 671B total parameters with 37B activated. Key improvements include smarter tool calling through post-training optimization, higher thinking efficiency achieving comparable quality to DeepSeek-R1-0528 while responding more quickly, and UE8M0 FP8 scale data format for model weights and activations. The model excels in both reasoning tasks (thinking mode) and practical applications (non-thinking mode), with particularly strong performance in code agent tasks, math competitions, and search-based problem solving.
Benchmark scores
| Benchmark | Score |
|---|---|
| SimpleQA | 0.93 |
| MMLU-Redux | 0.92 |
| MMLU-Pro | 0.84 |
| GPQA | 0.75 |
| AIME 2024 | 0.66 |
| SWE-Bench Verified | 0.66 |
| LiveCodeBench | 0.56 |
| AIME 2025 | 0.50 |
| Humanity's Last Exam | 0.16 |
Pricing
- DeepInfra$0.25 / $0.95
Input / output per 1M tokens
AA metrics
No Artificial Analysis link yet.
Arena Elo
- No Arena snapshot linked.
