Overview
HiDream-O1-Image-Pro is an image large model released by Zhipu Future on May 19, 2026, at the first Technology Open Day "Imaging the World". Built on the next-generation **native full-modal model architecture Unified Transformer (UiT)**, with over **200 billion** parameters, it **sets new SOTA records** on multiple benchmarks.
This is a key step for Zhipu Future from "visual generation" to "world model"—by unifying image pixels, text tokens, and task conditions into a continuous shared token space, the model can handle tasks such as general text-to-image, high-fidelity text rendering, and image editing in a unified manner. At the same time, Zhipu Future announced a new round of billion-level financing (with participation from Shenzhen Capital, Jinpu Investment, Caixin Capital, Fuju Capital, etc.), two rounds within half a month, indicating high recognition from the capital market for the native full-modal direction.
Key Features
- Over 200 Billion Parameters: Parameter scale reaches the leading level of domestic image large models, with model capacity ensuring generation quality
- Unified Transformer (UiT) Architecture: Native full-modal architecture, unifying image pixels, text tokens, and task conditions into a shared token space
- Multi-Benchmark SOTA: Refreshes industry best on multiple benchmarks including general text-to-image, high-fidelity text rendering, and image editing
- High-Fidelity Text Rendering: Industry-leading clarity and accuracy of Chinese and English text in images, suitable for posters, covers, and infographics
- Powerful Image Editing: Supports object replacement, style transfer, local modification, and other image editing capabilities
- Full-Modal Extensibility: Prepared for unified modeling of images, video, text, and audio, exploring the direction of world models
- Open-Source Version Available: The same architecture's 8B open-source version, HiDream-O1-Image, performs well on the Artificial Analysis text-to-image leaderboard
- Three Intelligent Agents Launched: Simultaneously released three intelligent agent products at the Open Day, transforming large model capabilities into applications
Use Cases
- Poster, cover, and infographic design requiring high-fidelity Chinese text rendering
- Batch generation of e-commerce product images and marketing materials
- Professional scenarios such as image editing, object replacement, and style transfer
- Enterprise users with needs for domestic AI images
- Academic research and prototype exploration (open-source version can be deployed locally)
- Content creators seeking domestic alternatives to Nano Banana / GPT-Image-2
Pros
- New benchmark for domestic AI image large models (200 billion parameters)
- Native full-modal architecture with leading technical path (not just a dedicated "text-to-image" model)
- Multi-benchmark SOTA with third-party validation of results
- Strong high-fidelity text rendering capability, especially suitable for Chinese scenarios
- 8B open-source version available for commercial use, lowering the barrier for domestic AI deployment
- Fast company financing pace (two rounds in half a month), with both technology and capital driving growth
- Evolving towards "world model", with significant future potential
Pricing
The Pro version is available through Zhipu Future's official platform and API, with specific pricing subject to official announcements. The same architecture's 8B open-source version, HiDream-O1-Image, has been open-sourced and can be downloaded for free for local deployment. Enterprise-level cooperation and private deployment are subject to separate business negotiations.
Summary
HiDream-O1-Image-Pro represents one of the highest levels of domestic AI image generation in 2026. In a track dominated by overseas models like Nano Banana / GPT-Image-2 / Imagen 4 / Midjourney, domestic models have found a differentiated breakthrough with "native full-modal architecture + 200 billion parameters + multi-benchmark SOTA". Zhipu Future's technical path from "visual generation" to "world model" gives it a unique strategic position in the domestic AI image direction. For users pursuing the latest SOTA results, needing high-fidelity Chinese rendering, and hoping to use domestic models, HiDream-O1-Image-Pro is one of the most noteworthy new products.