all the models — AI benchmark observatory
← Models

Llama 3.1 8B Instruct

Meta · open weight · llama-3.1-8b-instruct

Llama 3.1 8B Instruct is a multilingual large language model optimized for dialogue use cases. It features a 128K context length, state-of-the-art tool use, and strong reasoning capabilities.

IFEvalBFCLHumanEvalMMLUMMLU-ProGPQA

Benchmark scores

BenchmarkScore
IFEval0.80
BFCL0.76
HumanEval0.73
MMLU0.69
MMLU-Pro0.48
GPQA0.30

Pricing

  • DeepInfra$0.02 / $0.04

Input / output per 1M tokens

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.
Llama 3.1 8B Instruct Benchmarks · all the models