all the models — AI benchmark observatory
← Models

MAI-Thinking-1

Microsoft · proprietary · mai-thinking-1

MAI-Thinking-1 is Microsoft AI's first in-house reasoning model, a 35B-active / ~1T-total parameter sparse Mixture of Experts model (base model MAI-Base-1) trained from scratch without distillation from third-party models. Built with Microsoft's Hill-Climbing Machine pipeline, it was pre-trained on 30T tokens of clean, commercially licensed, human-generated data (plus 3.55T mid-training tokens), then post-trained via reinforcement learning across STEM, agentic coding, and helpfulness/safety specialists consolidated into a single model. It delivers strong mathematical reasoning and software-engineering performance for its weight class, going toe-to-toe with Claude Opus 4.6 on SWE-Bench Pro and reaching 97.0% on AIME 2025. It supports a 256k token context window, function calling, and developer instructions, and is preferred over Claude Sonnet 4.6 in blind human side-by-side evaluations.

AIME 2025AIME 2026MMLU-ProGPQASWE-Bench VeBFCL-v3

Benchmark scores

BenchmarkScore
AIME 20250.97
AIME 20260.94
MMLU-Pro0.85
GPQA0.84
SWE-Bench Verified0.73
BFCL-v30.72

Pricing

  • No provider pricing.

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.
MAI-Thinking-1 Benchmarks · all the models