← Models
MiMo-V2-Omni
Xiaomi · proprietary · mimo-v2-omni
MiMo-V2-Omni is Xiaomi's omni foundation model uniting frontier multimodal understanding with strong agentic capability. It fuses dedicated image, video, and audio encoders into a single shared backbone, processing all modalities simultaneously. Natively supports structured tool calling, function execution, and UI grounding. Supports over 10 hours of continuous audio understanding and 256K token context window.
Not enough scores for a fingerprint yet.
Benchmark scores
| Benchmark | Score |
|---|---|
| PinchBench | 0.81 |
| SWE-Bench Verified | 0.75 |
Pricing
- No provider pricing.
AA metrics
No Artificial Analysis link yet.
Arena Elo
- No Arena snapshot linked.
