all the models — AI benchmark observatory
← Models

Qwen2.5 14B Instruct

Alibaba Cloud / Qwen Team · open weight · qwen-2.5-14b-instruct

Qwen2.5-14B-Instruct is an instruction-tuned 14.7B parameter language model, part of the Qwen2.5 series. It features significant improvements in instruction following, long text generation (8K+ tokens), structured data understanding, and JSON output generation. The model supports a 128K token context length and multilingual capabilities across 29+ languages including Chinese, English, French, Spanish, and more.

GSM8kHumanEvalMATHMMLUMMLU-ProGPQA

Benchmark scores

BenchmarkScore
GSM8k0.95
HumanEval0.83
MATH0.80
MMLU0.80
MMLU-Pro0.64
GPQA0.46

Pricing

  • No provider pricing.

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.
Qwen2.5 14B Instruct Benchmarks · all the models