← Models
MiMo-V2.5-Pro
Xiaomi · open weight · mimo-v2.5-pro
MiMo-V2.5-Pro is Xiaomi's 1.02T-parameter sparse Mixture-of-Experts language model with 42B active parameters and a 1M-token context window. It inherits the MiMo-V2-Flash hybrid-attention and Multi-Token Prediction design, extends context during pre-training up to 1M tokens, and uses supervised fine-tuning, domain-specialized reinforcement learning, and Multi-Teacher On-Policy Distillation to improve complex software engineering, long-horizon agentic tasks, and ultra-long-context coherence.
Benchmark scores
| Benchmark | Score |
|---|---|
| FrontierSWE (Impl.) | 3.40 |
| GSM8k | 1.00 |
| ARC-C | 0.97 |
| MMLU-Redux | 0.93 |
| MMLU | 0.89 |
| MATH | 0.86 |
| SWE-Bench Verified | 0.79 |
| MiMo Coding Bench | 0.74 |
| TAU3-Bench | 0.73 |
| MMLU-Pro | 0.69 |
| GPQA | 0.67 |
| Humanity's Last Exam | 0.34 |
Pricing
- Xiaomi$0.44 / $0.87
- DeepInfra$1.00 / $3.00
- Novita$2.00 / $6.00
Input / output per 1M tokens
AA metrics
No Artificial Analysis link yet.
Arena Elo
- code1475
- text1468
