all the models — AI benchmark observatory
← Models

DeepSeek-V3 0324

DeepSeek · open weight · deepseek-v3-0324

A powerful Mixture-of-Experts (MoE) language model with 671B total parameters (37B activated per token). Features Multi-head Latent Attention (MLA), auxiliary-loss-free load balancing, and multi-token prediction training. Pre-trained on 14.8T tokens with strong performance in reasoning, math, and code tasks.

MATH-500MMLU-ProGPQAAIME 2024LiveCodeBenc

Benchmark scores

BenchmarkScore
MATH-5000.94
MMLU-Pro0.81
GPQA0.68
AIME 20240.59
LiveCodeBench0.49

Pricing

  • DeepInfra$0.24 / $0.90

Input / output per 1M tokens

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.