⚛️Stalecollected in 2h

HiDream-O1-Image-1.5 Tops Global Text-to-Image Benchmarks

HiDream-O1-Image-1.5 Tops Global Text-to-Image Benchmarks
PostLinkedIn
⚛️Read original on 量子位

💡A new Chinese model is outperforming Google and Nvidia on global text-to-image benchmarks.

⚡ 30-Second TL;DR

What Changed

HiDream-O1-Image-1.5 ranks first in China and second globally in image generation.

Why It Matters

This breakthrough signals increased competitiveness of Chinese generative models on the global stage. It challenges the dominance of Western tech giants in high-fidelity image synthesis.

What To Do Next

Evaluate HiDream-O1-Image-1.5 against your current image generation pipeline to assess potential quality improvements.

Who should care:Researchers & Academics

Key Points

  • HiDream-O1-Image-1.5 ranks first in China and second globally in image generation.
  • The model demonstrates superior performance compared to Google and Nvidia offerings.
  • This marks a significant milestone for Chinese AI research in generative visual models.

🧠 Deep Insight

Web-grounded analysis with 19 cited sources.

🔑 Enhanced Key Takeaways

  • HiDream.ai, the company behind the model, was founded in Beijing, China, in 2023 by Tao Mei, who is a foreign academician of the Canadian Academy of Engineering and former vice president of Jingdong Group.
  • The HiDream-O1-Image-1.5 is a commercial version, while an open-source 8-billion-parameter model, HiDream-O1-Image-Dev, is also available under an MIT license and has shown performance comparable to or surpassing larger open-source models.
  • The model employs a novel Pixel-level Unified Transformer (UiT) architecture that operates directly on raw pixels, eliminating the need for a separate Variational Autoencoder (VAE) or disjoint text encoders, and natively supports high-resolution output up to 2048x2048.
  • HiDream.ai recently secured a Series C funding round in May 2026, with investments from Shenzhen Capital Real Estate Fund, GP Capital, Caixin Capital, and Zhejiang Fuju Investment Management Co., Ltd., indicating strong investor confidence.
  • Beyond text-to-image generation, HiDream-O1-Image-1.5 offers multimodal capabilities including instruction-based image editing, subject-driven personalization, accurate long-text rendering, and storyboard generation within a single architecture.
📊 Competitor Analysis▸ Show
ModelDeveloperArchitectureParametersKey FeaturesBenchmarks (ELO Score - Source)Pricing (per 1k images)
HiDream-O1-Image-1.5 (Commercial)HiDream.aiPixel-level Unified Transformer (UiT)200B+ (Pro version)Text-to-image, editing, personalization, long-text rendering, storyboard, native 2K resolution, Reasoning-Driven Prompt AgentRanks 1st in China, 2nd globally (量子位)N/A (Commercial, likely enterprise)
HiDream-O1-Image-Dev (Open-source)HiDream.aiPixel-level Unified Transformer (UiT)8BText-to-image, editing, personalization, long-text rendering, storyboard, native 2K resolution, Reasoning-Driven Prompt Agent1192 ELO (Artificial Analysis)N/A (Open-source, MIT license)
GPT Image 2 (high)OpenAIProprietaryN/AHigh fidelity, best pose/anatomical accuracy, best for mockups1340 ELO (Artificial Analysis)$211.0
Nano Banana 2 (Gemini 3.1 Flash Image Preview)GoogleProprietaryN/AImage generation, editing, multi-turn editing, 14 reference images, 4K resolution, Google Search grounding1260 ELO (Artificial Analysis)$67.0
Cosmos3-Super-Text2Image (agentic)NVIDIAMixture-of-Transformers (Omnimodel)64BPhysical AI, vision reasoning, multimodal generation (text, image, video, sound, action), synthetic data generation1239 ELO (Artificial Analysis)Coming soon
Ideogram 4.0IdeogramProprietary (Open-weight)N/ATop-ranked open-weight, accurate text rendering, LoRA fine-tuningTop spot for open-weight (Artificial Analysis)N/A (API available)

🛠️ Technical Deep Dive

  • HiDream-O1-Image is built on a Pixel-level Unified Transformer (UiT) architecture.
  • It operates directly on raw pixels, eliminating the need for an external Variational Autoencoder (VAE) or disjoint text encoders.
  • The model uses a single shared token space for raw pixels, text prompts, and task-specific conditions.
  • The open-source version, HiDream-O1-Image (and its Dev variant), has 8 billion parameters.
  • The commercial 'Pro' version (HiDream-O1-Image-Pro) features over 200 billion parameters.
  • It supports native high-resolution image synthesis up to 2048x2048 pixels.
  • The model incorporates a 'Reasoning-Driven Prompt Agent' that resolves implicit knowledge, layout, and text rendering before generation.
  • It is designed for multiple tasks, including text-to-image generation, instruction-based image editing, subject-driven personalization, long-text rendering, and storyboard generation.
  • The distilled Dev variant can converge in 28 inference steps with a guidance scale of 0, while the full model uses 50 steps with a CFG of 5.0.

🔮 Future ImplicationsAI analysis grounded in cited sources

The novel pixel-level architecture could set a new industry standard.
By eliminating VAEs and operating directly on raw pixels, HiDream-O1-Image's architecture addresses common challenges in fine detail and text rendering, potentially influencing future model designs across the industry.
HiDream.ai is poised to strengthen China's position in the global generative AI landscape.
The model's top-tier performance on global benchmarks, combined with significant Series C funding, indicates HiDream.ai's growing influence and competitive edge in advanced AI research and commercialization.
The multimodal capabilities suggest a broader strategic vision beyond static image generation.
With support for image, video, and 3D content generation, HiDream.ai is positioning itself as a comprehensive multimodal content generation and application developer, aiming for a more complete 'world modeling' ability.

Timeline

2023
HiDream.ai founded by Tao Mei in Beijing, China
2025-03
Original HiDream-I1, a 17B sparse-MoE DiT model, was released
2026-05-08
HiDream-O1-Image (8B parameters) and its distilled Dev variant, along with the Reasoning-Driven Prompt Agent, were open-sourced under an MIT license
2026-05-10
The technical report for HiDream-O1-Image was made available
2026-05-19
HiDream.ai announced the completion of a Series C funding round
2026-06-10
HiDream-O1-Image-1.5 was reported to rank first in China and second globally on text-to-image benchmarks, surpassing Google and Nvidia
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: 量子位