← Models
DeepSeek-V3
DeepSeek · open weight · deepseek-v3
A powerful Mixture-of-Experts (MoE) language model with 671B total parameters (37B activated per token). Features Multi-head Latent Attention (MLA), auxiliary-loss-free load balancing, and multi-token prediction training. Pre-trained on 14.8T tokens with strong performance in reasoning, math, and code tasks.
Benchmark scores
| Benchmark | Score |
|---|---|
| DROP | 0.92 |
| MATH-500 | 0.90 |
| MMLU-Redux | 0.89 |
| MMLU | 0.89 |
| IFEval | 0.86 |
| Aider-Polyglot Edit | 0.80 |
| MMLU-Pro | 0.76 |
| GPQA | 0.59 |
| SWE-Bench Verified | 0.42 |
| AIME 2024 | 0.39 |
| LiveCodeBench | 0.38 |
Pricing
- DeepInfra$0.32 / $0.89
Input / output per 1M tokens
AA metrics
No Artificial Analysis link yet.
Arena Elo
- No Arena snapshot linked.
