all the models — AI benchmark observatory
← Models

Phi 4

Microsoft · open weight · phi-4

phi-4 is a state-of-the-art open model built to excel at advanced reasoning, coding, and knowledge tasks. It leverages a blend of synthetic data, filtered web data, academic texts, and supervised fine-tuning for precision, alignment, and safety.

MMLUHumanEvalMATHArena HardMMLU-ProIFEval

Benchmark scores

BenchmarkScore
MMLU0.85
HumanEval0.83
MATH0.80
Arena Hard0.75
MMLU-Pro0.70
IFEval0.63
GPQA0.56

Pricing

  • DeepInfra$0.07 / $0.14

Input / output per 1M tokens

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.