Overview
MiMo-V2.6-Pro-UltraSpeed is the high-speed serving tier Xiaomi added to the MiMo line on September 22, 2026, alongside the V2.6 release. It runs the same checkpoint as V2.6-Pro: 1.0T total parameters, 42B activated, a 1M-token context, and text, image, video and audio input. The capability set is identical, and what changes is the inference path. Xiaomi advertises up to 20x output speed, trading price for time.
Pricing is the clearest part of the offer. Per million tokens, domestic rates run 0.25 yuan on cache hits, 30 yuan for input and 60 yuan for output; overseas the same three come to 0.036, 4.35 and 8.7 dollars, roughly ten times the standard Pro tier. Third-party coverage consistently describes it as about ten times the speed for about ten times the price, and the 20x claim carries no independent measurement yet. Before moving a workload across, run your own tasks and read latency and quality together.
There are two ways in. The open platform exposes UltraSpeed as an option under the V2.6-Pro endpoint, and MiMo Desktop lists it on the top two membership tiers. Batch inference covers Pro and Flash; UltraSpeed bills through the real-time path, with cache writes free for a limited time.
Those numbers also draw the boundary. It fits real-time interaction where the rhythm has to match a human, voice and conversational products that treat first-token latency as a product metric, and production pipelines already on V2.6-Pro where response time has become the bottleneck. Long-form writing and bulk data processing that are not on a clock stay cheaper on the standard tier.
Key Features
- Same checkpoint as the standard tier: The capability set matches V2.6-Pro, so migration means switching a model name rather than rebuilding prompts, tooling or evaluation baselines.
- Up to 20x output speed, per Xiaomi: Aimed at real-time interaction and latency-sensitive production work; third-party coverage mostly reports about ten times the actual speed, and independent measurement has yet to appear.
- Trillion-parameter omni-modal base: A sparse MoE with 1.0T total and 42B active parameters, 1M-token context, and text, image, video and audio input with text output.
- Layered pricing that is easy to read: Domestically 0.25 yuan on cache hits, 30 yuan input and 60 yuan output per million tokens, ten times the unit price for the speed, so splitting a workload by task is straightforward.
- Available on both desktop and platform: The open platform offers the option under the V2.6-Pro endpoint, while MiMo Desktop marks it as exclusive to the top two membership tiers.
- Separate real-time billing: Batch inference covers Pro and Flash, while UltraSpeed bills through the real-time path, with cache writes currently free for a limited time.
Use Cases
- Voice conversation and live interpretation, where character-by-character output has to keep pace with human speech
- IDE inline completion and long agent chains, the interaction-heavy spots where waiting is most noticeable
- Support agents, live-broadcast assistance and game NPC dialogue that need immediate feedback
- Production workflows already on V2.6-Pro where response time has become the bottleneck, switching only the critical path
Pros
- Identical checkpoint to V2.6-Pro, so capability holds and switching cost stays low
- The jump in output speed changes interactive experience by an order of magnitude
- Transparent domestic pricing, ten times the cost for the speed, making tier choice a clear call
- Reachable from both the open platform and the desktop client
Pricing
Per million tokens domestically: 0.25 yuan on cache hits, 30 yuan input, 60 yuan output. Overseas: 0.036, 4.35 and 8.7 dollars. Batch inference covers Pro and Flash, while UltraSpeed bills through the real-time path with cache writes free for a limited time. MiMo Desktop marks the model as exclusive to its top two membership tiers. Check official pricing pages for current quotas and rates.
Summary
MiMo-V2.6-Pro-UltraSpeed is the high-speed serving tier Xiaomi built for V2.6-Pro: the same checkpoint and the same omni-modal capability, with the inference path pushing output speed to a claimed 20x at ten times the price. The gain in interactive feel is real, and so is the change in cost structure, which makes it a fit for critical paths rather than wholesale replacement. The test is simple: if waiting itself is hurting the product experience or developer throughput, this tier is worth trying; if the work is not on a clock, standard Pro stays cheaper.
Version History
- MiMo-V2.6-Pro-UltraSpeed ships with the V2.6 family, advertising up to 20x output speed (2026-09-22): Xiaomi launched a high-speed serving tier for Pro alongside the MiMo-V2.6 family, sharing the 1.0T total and 42B active checkpoint plus the 1M-token context, with up to 20x output speed as the advertised figure. Domestic pricing per million tokens is 0.25 yuan on cache hits, 30 yuan input and 60 yuan output; overseas it is 0.036, 4.35 and 8.7 dollars, roughly ten times the standard tier, which Xiaomi frames as ten times the calling cost for up to 20x output speed. The open platform offers it under the V2.6-Pro endpoint, MiMo Desktop marks it exclusive to the top two membership tiers, and batch inference covers Pro and Flash.