all the models — AI benchmark observatory
← Models

MiMo-V2-Omni

Xiaomi · proprietary · mimo-v2-omni

MiMo-V2-Omni is Xiaomi's omni foundation model uniting frontier multimodal understanding with strong agentic capability. It fuses dedicated image, video, and audio encoders into a single shared backbone, processing all modalities simultaneously. Natively supports structured tool calling, function execution, and UI grounding. Supports over 10 hours of continuous audio understanding and 256K token context window.

Not enough scores for a fingerprint yet.

Benchmark scores

BenchmarkScore
PinchBench0.81
SWE-Bench Verified0.75

Pricing

  • No provider pricing.

AA metrics

No Artificial Analysis link yet.

Arena Elo

  • No Arena snapshot linked.
MiMo-V2-Omni Benchmarks · all the models