HiDream-O1-Image-1.5 Tops Global Text-to-Image Benchmarks

💡A new Chinese model is outperforming Google and Nvidia on global text-to-image benchmarks.
⚡ 30-Second TL;DR
What Changed
HiDream-O1-Image-1.5 ranks first in China and second globally in image generation.
Why It Matters
This breakthrough signals increased competitiveness of Chinese generative models on the global stage. It challenges the dominance of Western tech giants in high-fidelity image synthesis.
What To Do Next
Evaluate HiDream-O1-Image-1.5 against your current image generation pipeline to assess potential quality improvements.
Key Points
- •HiDream-O1-Image-1.5 ranks first in China and second globally in image generation.
- •The model demonstrates superior performance compared to Google and Nvidia offerings.
- •This marks a significant milestone for Chinese AI research in generative visual models.
🧠 Deep Insight
Web-grounded analysis with 19 cited sources.
🔑 Enhanced Key Takeaways
- •HiDream.ai, the company behind the model, was founded in Beijing, China, in 2023 by Tao Mei, who is a foreign academician of the Canadian Academy of Engineering and former vice president of Jingdong Group.
- •The HiDream-O1-Image-1.5 is a commercial version, while an open-source 8-billion-parameter model, HiDream-O1-Image-Dev, is also available under an MIT license and has shown performance comparable to or surpassing larger open-source models.
- •The model employs a novel Pixel-level Unified Transformer (UiT) architecture that operates directly on raw pixels, eliminating the need for a separate Variational Autoencoder (VAE) or disjoint text encoders, and natively supports high-resolution output up to 2048x2048.
- •HiDream.ai recently secured a Series C funding round in May 2026, with investments from Shenzhen Capital Real Estate Fund, GP Capital, Caixin Capital, and Zhejiang Fuju Investment Management Co., Ltd., indicating strong investor confidence.
- •Beyond text-to-image generation, HiDream-O1-Image-1.5 offers multimodal capabilities including instruction-based image editing, subject-driven personalization, accurate long-text rendering, and storyboard generation within a single architecture.
📊 Competitor Analysis▸ Show
| Model | Developer | Architecture | Parameters | Key Features | Benchmarks (ELO Score - Source) | Pricing (per 1k images) |
|---|---|---|---|---|---|---|
| HiDream-O1-Image-1.5 (Commercial) | HiDream.ai | Pixel-level Unified Transformer (UiT) | 200B+ (Pro version) | Text-to-image, editing, personalization, long-text rendering, storyboard, native 2K resolution, Reasoning-Driven Prompt Agent | Ranks 1st in China, 2nd globally (量子位) | N/A (Commercial, likely enterprise) |
| HiDream-O1-Image-Dev (Open-source) | HiDream.ai | Pixel-level Unified Transformer (UiT) | 8B | Text-to-image, editing, personalization, long-text rendering, storyboard, native 2K resolution, Reasoning-Driven Prompt Agent | 1192 ELO (Artificial Analysis) | N/A (Open-source, MIT license) |
| GPT Image 2 (high) | OpenAI | Proprietary | N/A | High fidelity, best pose/anatomical accuracy, best for mockups | 1340 ELO (Artificial Analysis) | $211.0 |
| Nano Banana 2 (Gemini 3.1 Flash Image Preview) | Proprietary | N/A | Image generation, editing, multi-turn editing, 14 reference images, 4K resolution, Google Search grounding | 1260 ELO (Artificial Analysis) | $67.0 | |
| Cosmos3-Super-Text2Image (agentic) | NVIDIA | Mixture-of-Transformers (Omnimodel) | 64B | Physical AI, vision reasoning, multimodal generation (text, image, video, sound, action), synthetic data generation | 1239 ELO (Artificial Analysis) | Coming soon |
| Ideogram 4.0 | Ideogram | Proprietary (Open-weight) | N/A | Top-ranked open-weight, accurate text rendering, LoRA fine-tuning | Top spot for open-weight (Artificial Analysis) | N/A (API available) |
🛠️ Technical Deep Dive
- HiDream-O1-Image is built on a Pixel-level Unified Transformer (UiT) architecture.
- It operates directly on raw pixels, eliminating the need for an external Variational Autoencoder (VAE) or disjoint text encoders.
- The model uses a single shared token space for raw pixels, text prompts, and task-specific conditions.
- The open-source version, HiDream-O1-Image (and its Dev variant), has 8 billion parameters.
- The commercial 'Pro' version (HiDream-O1-Image-Pro) features over 200 billion parameters.
- It supports native high-resolution image synthesis up to 2048x2048 pixels.
- The model incorporates a 'Reasoning-Driven Prompt Agent' that resolves implicit knowledge, layout, and text rendering before generation.
- It is designed for multiple tasks, including text-to-image generation, instruction-based image editing, subject-driven personalization, long-text rendering, and storyboard generation.
- The distilled Dev variant can converge in 28 inference steps with a guidance scale of 0, while the full model uses 50 steps with a CFG of 5.0.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (19)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates

Meta testing StoryKit for AI-generated children's stories
Alibaba Releases Qwen-Image-3.0 Generation Model

Shanghai AI competition focuses on autonomous research and fusion

Shanghai IEDG Establishes Intelligent Computing System Architecture Alliance
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 量子位 ↗