Overview
MiMo-V2.5-TTS is an **end-to-end speech synthesis large model officially released by Xiaomi on April 24, 2026**, part of the MiMo-V2.5 full-chain speech model series (TTS synthesis + ASR recognition dual capabilities) — this marks Xiaomi's formal deployment of the complete speech input-output chain in the Agent era.
**Core selling points of MiMo-V2.5-TTS**:
- **Natural language voice control**: Unlike traditional TTS requiring complex parameter configurations (pitch, speed, emotion intensity, pauses...), MiMo-V2.5-TTS can be **directly controlled via natural language commands** — for example, "use a gentle tone of a 30-year-old woman, slightly faster, with a hint of surprise," and the model understands and outputs the corresponding effect. This is the first time in the TTS field to truly achieve "director-style control."
- **Native dialect support**: Major dialects such as Cantonese, Sichuanese, Henan dialect, and Taiwanese accent are output natively, not simple accent imitation but real dialect grammar and pronunciation (a major breakthrough for content creation, short videos, and audiobooks).
- **Emotion switching and singing capability**: Emotion switching within the same text (e.g., "pretend to be angry here, then pretend to cry"), supports singing pitch control (the model can sing, not just speak).
- **Limited-time free API**: TTS API is currently free for a limited time, and ASR is fully open source (`mimo.xiaomi.com`).
**Three sub-models**:
- **Basic version**: Fine control over multiple timbres (preset timbres + parameter fine-tuning).
- **VoiceDesign version**: Generate new timbres through description (describe the desired voice in text).
- **Voice cloning version**: Clone any voice with one sentence (3 seconds).
**Product positioning**: Aimed at speech output needs in the AI Agent era — Agents no longer only reply with fixed synthetic voices but **dynamically select appropriate voice expressions based on the scenario**. Paired with the MiMo-V2.5-ASR speech recognition model, it forms a complete speech IO loop.
It is one of the most noteworthy open-source/free products in the domestic TTS field in 2026.
Key Features
- Natural language voice control (pioneering): Adjust timbre/emotion/speed/pauses by 'talking like a director,' no complex parameters needed.
- Native dialect support: Real grammar and pronunciation for Cantonese, Sichuanese, Henan dialect, and Taiwanese accent (not accent imitation).
- Emotion switching: Emotion changes within the same text (angry → cry → laugh), matching human expression.
- Singing pitch control: The model can sing (not speak), with fine pitch control.
- One-sentence voice cloning: Clone any voice with 3 seconds of audio (VoiceDesign version).
- Voice design version: Generate entirely new timbres that never existed through text description (VoiceDesign).
- TTS limited-time free + ASR fully open source: TTS API is currently free for a limited time; ASR speech recognition model is fully open source under MIT license.
Use Cases
- Audiobook / podcast production (multi-role + emotion switching)
- Short video dubbing (dialect + emotion + singing)
- AI digital human speech output (director-style control)
- Game / anime character dubbing
- Accessibility assistance (natural speech output for visually impaired users)
- Dynamic voice selection for Agent voice assistants
- Diverse voice content in education/training scenarios
Pros
- Natural language voice control is an industry first, offering a generational leap in user experience.
- Native dialect output (Cantonese/Sichuanese/Henan dialect/Taiwanese accent) fills a market gap.
- Emotion switching + singing capability covers all content creation scenarios.
- One-sentence voice cloning has an extremely low barrier (3 seconds).
- TTS limited-time free reduces trial cost, ASR fully open source.
- Backed by Xiaomi, ensuring model capability and stability.
- Designed for the Agent era, deeply integrated with the speech IO ecosystem.
Pricing
**TTS API is limited-time free** (apply directly on the official website `mimo.xiaomi.com`): All three sub-models (Basic, VoiceDesign, Voice Cloning) can be called for free, with specific quotas subject to official announcements. **ASR speech recognition model is fully open source** (MIT license, weights downloadable on Hugging Face / GitHub). **Commercial pricing plans** will be introduced in the future, differentiating between individual developers and enterprise customers.
Summary
MiMo-V2.5-TTS is a phenomenal product in the domestic TTS field in 2026 — the combination of **natural language voice control + native dialects + emotion switching + singing capability + limited-time free** makes it the top choice for content creators, AI digital human developers, and Agent voice product teams. **If you are working on short video dubbing, audiobooks, podcasts, or game dubbing**, MiMo-V2.5-TTS is a must-test — especially for dialect scenarios, it has almost no domestic competitors. Paired with MiMo-V2.5-ASR, it forms a complete speech IO loop, representing Xiaomi's comprehensive deployment of the full speech chain in the Agent era.
Version History
- MiMo-V2.5 full-chain speech series release (2026-04-24): TTS (three sub-models: Basic + VoiceDesign + Voice Cloning) + ASR dual capabilities released; natural language voice control industry first; native dialect support; TTS limited-time free + ASR fully open source