all the models — AI benchmark observatory
← Models

Phi 4 Reasoning Plus

Microsoft · open weight · phi-4-reasoning-plus

Phi-4-reasoning-plus is a state-of-the-art open-weight reasoning model finetuned from Phi-4 using supervised fine-tuning and reinforcement learning. It focuses on math, science, and coding skills. This 'plus' version has higher accuracy due to additional RL training but may have higher latency.

FlenQAHumanEval+IFEvalAIME 2024Arena HardAIME 2025

Benchmark scores

BenchmarkScore
FlenQA0.98
HumanEval+0.92
IFEval0.85
AIME 20240.81
Arena Hard0.79
AIME 20250.78
MMLU-Pro0.76
GPQA0.69
LiveCodeBench0.53

Pricing

  • No provider pricing.

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.
Phi 4 Reasoning Plus Benchmarks · all the models