Overview
MiMo-V2.5-Pro UltraSpeed is an ultra-fast inference version jointly released by Xiaomi and **TileRT**, **the world's first model to achieve a decoding speed of 1000 tokens/s at the 1T parameter scale** (peak 1200 t/s). Its product philosophy is '**trade raw speed for depth of thought**' — on the capability base of a 1 trillion parameter MoE, it compresses inference latency to a level where 'the human eye can't keep up with the output,' enabling real-time follow-up of the AI's thinking process in coding sessions instead of 'waiting tens of seconds for a paragraph.'
Behind this is TileRT's self-developed inference acceleration stack: a combination of multiple engineering techniques including **operator fusion + speculative decoding + speculative sampling + quantization symbiosis**. The most intuitive impact for developers is a qualitative change in the Coding Agent experience — previously, waiting 30 seconds for AI to think before outputting code in Cursor / Claude Code, now UltraSpeed turns waiting into 'instant screen refresh right after asking,' reducing the total time for long conversations and long-range programming tasks with 200+ steps to less than one-third.
This is one of the representative achievements of LLM engineering in 2026: **once model capabilities hit a ceiling, the next wave of competition is inference infrastructure.**
Key Features
- 1000 t/s Decoding Speed: The world's first model to achieve 1000 t/s decoding speed at the 1T parameter scale, with a peak of 1200 t/s, far exceeding the typical 30-100 t/s of mainstream frontier models.
- 1T Parameter MoE Base: The base is the MiMo-V2.5-Pro trillion-parameter MoE model, with capabilities comparable to Claude Opus 4.6; UltraSpeed does not sacrifice underlying capability for speed.
- TileRT Self-Developed Inference Stack: A combination of multiple engineering techniques including operator fusion, speculative decoding, speculative sampling, and quantization symbiosis, pushing inference latency to the extreme.
- Real-Time Follow-Up in Coding Agent: In Coding Agent scenarios, the AI's thinking process can be followed in real time, reducing total time for long conversations and long-range programming tasks with 200+ steps to less than one-third.
- Available on MiMo Open Platform: Directly callable through Xiaomi's MiMo Open Platform, using the same API as MiMo-V2.5-Pro, allowing smooth switching.
Use Cases
- High-frequency Coding Agent scenarios (200+ step long-range programming, long conversations)
- Real-time assistance sensitive to inference latency (IDE inline completion, real-time translation, conversational BI)
- Large-scale batch processing tasks (document processing, data cleaning, content generation pipelines) requiring cost reduction and efficiency improvement
- Products like AI customer service / virtual assistants that need a 'seamless waiting' experience
- Heavy users of MiMo-V2.5-Pro upgrading to UltraSpeed for efficiency gains
Pros
- 1000 t/s is a global first at the 1T parameter scale, with significant engineering implications.
- Base capability is uncompromised (still MiMo-V2.5-Pro trillion MoE).
- Qualitative change in Coding Agent experience.
- TileRT's self-developed inference stack is significant for domestic inference infrastructure.
- Direct connection to MiMo Open Platform with low migration cost.
Pricing
Called through Xiaomi's MiMo Open Platform, pricing is differentiated from the standard MiMo-V2.5-Pro API (speed tier is more expensive than standard tier). After the MiMo platform's permanent 99% price reduction on 2025/05/27, the overall unit price remains low in the industry. Specific unit prices are subject to announcements on platform.xiaomimimo.com.
Summary
MiMo-V2.5-Pro UltraSpeed represents a new direction for large model engineering in 2026 — **once model capabilities plateau, the next wave is inference infrastructure**. 1000 t/s is not just about being 'fast'; it is a paradigm shift in AI programming experience from 'wait → see → edit' to 'ask and get instantly.' If you heavily use Coding Agents, are sensitive to latency, or handle high-frequency batch processing or real-time assistant products, UltraSpeed is currently one of the best balances of 'capability + speed.' For everyday conversations, the standard MiMo-V2.5-Pro tier is fully sufficient.