← Models
DeepSeek-V3 0324
DeepSeek · open weight · deepseek-v3-0324
A powerful Mixture-of-Experts (MoE) language model with 671B total parameters (37B activated per token). Features Multi-head Latent Attention (MLA), auxiliary-loss-free load balancing, and multi-token prediction training. Pre-trained on 14.8T tokens with strong performance in reasoning, math, and code tasks.
Benchmark scores
| Benchmark | Score |
|---|---|
| MATH-500 | 0.94 |
| MMLU-Pro | 0.81 |
| GPQA | 0.68 |
| AIME 2024 | 0.59 |
| LiveCodeBench | 0.49 |
Pricing
- DeepInfra$0.24 / $0.90
Input / output per 1M tokens
AA metrics
No Artificial Analysis link yet.
Arena Elo
- No Arena snapshot linked.
