all the models — AI benchmark observatory
← Models

DeepSeek-V3

DeepSeek · open weight · deepseek-v3

A powerful Mixture-of-Experts (MoE) language model with 671B total parameters (37B activated per token). Features Multi-head Latent Attention (MLA), auxiliary-loss-free load balancing, and multi-token prediction training. Pre-trained on 14.8T tokens with strong performance in reasoning, math, and code tasks.

DROPMATH-500MMLU-ReduxMMLUIFEvalAider-Polygl

Benchmark scores

Pricing

  • DeepInfra$0.32 / $0.89

Input / output per 1M tokens

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.