all the models — AI benchmark observatory
← Models

Qwen3 Max Thinking

Alibaba Cloud / Qwen Team · proprietary · qwen3-max-thinking

The latest flagship reasoning model in the Qwen3 family. Further enhanced by multiple innovations like adaptive tool-use and advanced test-time scaling techniques

C-EvalIFEvalAIME 2026MMLU-ReduxGPQAMMLU-Pro

Benchmark scores

BenchmarkScore
C-Eval0.94
IFEval0.93
AIME 20260.93
MMLU-Redux0.93
GPQA0.87
MMLU-Pro0.86
SWE-Bench Verified0.75

Pricing

  • DeepInfra$1.20 / $6.00

Input / output per 1M tokens

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.