← Models
MiMo-V2.5
Xiaomi · open weight · mimo-v2.5
MiMo-V2.5 is Xiaomi's native omnimodal sparse Mixture-of-Experts model with 310B total parameters, 15B activated parameters, and a 1M-token context window. Built on the MiMo-V2-Flash backbone, it adds dedicated vision and audio encoders for text, image, video, and audio understanding, and is post-trained with SFT, agentic reinforcement learning, and Multi-Teacher On-Policy Distillation for multimodal perception, long-context reasoning, and agentic workflows.
Benchmark scores
| Benchmark | Score |
|---|---|
| HR-Bench (4k) | 0.89 |
| Video-MME | 0.88 |
| OmniDocBench | 0.87 |
| MMMU-Pro | 0.78 |
| MiMo Coding Bench | 0.72 |
Pricing
- Novita$0.17 / $0.34
- DeepInfra$0.40 / $2.00
Input / output per 1M tokens
AA metrics
No Artificial Analysis link yet.
Arena Elo
- No Arena snapshot linked.
