all the models — AI benchmark observatory
← Models

Llama 3.1 70B Instruct

Meta · open weight · llama-3.1-70b-instruct

Llama 3.1 70B Instruct is a large language model optimized for multilingual dialogue use cases. It outperforms many available open source and closed chat models on common industry benchmarks.

GSM-8K (CoT)ARC-CIFEvalBFCLMMLUHumanEval

Benchmark scores

BenchmarkScore
GSM-8K (CoT)0.95
ARC-C0.95
IFEval0.88
BFCL0.85
MMLU0.84
HumanEval0.81
MMLU-Pro0.66
GPQA0.42

Pricing

  • DeepInfra$0.40 / $0.40

Input / output per 1M tokens

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.
Llama 3.1 70B Instruct Benchmarks · all the models