all the models — AI benchmark observatory
← Models

GPT-5.1

OpenAI · proprietary · gpt-5.1-2025-11-13

The best model for coding and agentic tasks with configurable reasoning effort. GPT-5.1 is our flagship model for coding and agentic tasks with configurable reasoning and non-reasoning effort.

Tau2 TelecomAIME 2025GPQAMMMUSWE-Bench Ve

Benchmark scores

BenchmarkScore
Tau2 Telecom0.96
AIME 20250.94
GPQA0.88
MMMU0.85
SWE-Bench Verified0.76

Pricing

  • OpenAI$1.25 / $10.00

Input / output per 1M tokens

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.