all the models — AI benchmark observatory
← Models

Llama 3.3 70B Instruct

Meta · open weight · llama-3.3-70b-instruct

Llama 3.3 is a multilingual large language model optimized for dialogue use cases across multiple languages. It is a pretrained and instruction-tuned generative model with 70 billion parameters, outperforming many open-source and closed chat models on common industry benchmarks. Llama 3.3 supports a context length of 128,000 tokens and is designed for commercial and research use in multiple languages.

IFEvalMGSMHumanEvalMMLUMATHMMLU-Pro

Benchmark scores

BenchmarkScore
IFEval0.92
MGSM0.91
HumanEval0.88
MMLU0.86
MATH0.77
MMLU-Pro0.69
GPQA0.51

Pricing

  • DeepInfra$0.10 / $0.32

Input / output per 1M tokens

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.