Overview
MiMo-V2-Omni is a **full-modal base model** released by Xiaomi in 2026, serving as the flagship product of Xiaomi's AI strategy of "end-cloud synergy + full modality." "Omni" means **native unified processing of images, videos, audio, and text**—not by stitching together single-modal models, but by supporting any modality as input or output at the architectural level, including cross-modal reasoning (watching a video and answering a voice question, reading an image and generating related audio commentary, etc.).
Xiaomi's official **PinchBench** (multimodal comprehensive benchmark) scores **lead Gemini 3 Pro and Claude Opus 4.6** in multiple areas, especially in video understanding, long audio dialogue, and interleaved image-text reasoning, placing it in the top tier among current open-source/domestic models. The model's weights are open-sourced under the **MIT license**, making it one of the most generous releases among domestic full-modal open-source models.
Together with the same series' **MiMo-V2.5-Pro** (reasoning/coding flagship) and **MiMo-V2.5-TTS** (speech synthesis), it forms a complete product matrix: Omni handles cross-modal understanding and generation, Pro handles complex reasoning and coding, and TTS handles high-quality speech output. Together, they support the on-device and cloud AI experiences of Xiaomi phones, cars, and smart home devices.
Key Features
- Native Full-Modal Unified Architecture: Any input + any output for images/videos/audio/text, with cross-modal reasoning as a default capability rather than an add-on module
- Leading on PinchBench: Multiple scores on Xiaomi's official PinchBench multimodal benchmark lead Gemini 3 Pro and Claude Opus 4.6
- Video and Long Audio Understanding: Native support for understanding, summarizing, and Q&A on tens-of-minutes-long videos and audio, which is Omni's strongest area
- Interleaved Image-Text Reasoning: Handles interleaved input of "text + image + text + image" for complex reasoning, covering scenarios like papers, technical documents, and multi-image manuals
- Open Source under MIT License: Weights and inference code are open under the MIT license, allowing commercial use and fine-tuning, making it most developer-friendly
- End-Cloud Synergy Deployment: A distilled version can run on Xiaomi's HyperOS on-device, while the cloud version provides full capabilities, ensuring a consistent experience across Xiaomi's entire product line
Use Cases
- Full-modal assistant capabilities on Xiaomi phones/cars/smart home devices
- Video content understanding: automatic summarization, chapter segmentation, Q&A, keyframe localization
- Education scenarios: interleaved reasoning with images, text, and formulas, subject Q&A, video lecture explanations
- Industry/manufacturing: defect recognition directly from camera video or images, plus text report generation
- Accessibility: image description, video commentary, cross-modal dialogue, benefiting visually/hearing-impaired users
- Developers fine-tuning industry-specific full-modal models based on open-source weights
Pros
- Native full-modal architecture, with cross-modal task capabilities in the top tier among domestic models
- Multiple PinchBench scores lead Gemini 3 Pro and Claude Opus 4.6
- MIT license open-sourcing is one of the most generous release stances among major domestic companies
- Combined with MiMo-V2.5-Pro/TTS, covers reasoning/coding/full-modal/speech scenarios comprehensively
- End-cloud synergy: on-device distillation + full cloud version, amplifying value across Xiaomi's entire product matrix
Pricing
**Open-source weights are free** (MIT license, downloadable from HuggingFace/GitHub). **Official API** (mimo.xiaomi.com) offers a free trial quota, with overage billed per token, at a price range significantly lower than Gemini 3 Pro and Claude Opus 4.6. Xiaomi phone/car users can use full-modal capabilities **completely free** in system-integrated scenarios. Enterprise private deployment can be quoted per GPU node.
Summary
MiMo-V2-Omni is one of the most significant releases of domestic/open-source full-modal models in 2026—the combination of **native unified architecture + leading PinchBench scores + MIT open-source** gives it a clear differentiated position in the field. If you are building cross-modal applications (video Q&A, multimodal agents, comprehensive image-text-audio understanding), Omni is the top choice among current open-source models. If you need reasoning/coding, consider the same series' **MiMo-V2.5-Pro**; for speech synthesis, use **MiMo-V2.5-TTS**. Together, the trio forms the most complete product matrix for domestic open-source full-modal models.