來源較早收集於 2h

HiDream-O1-Image-1.5 登頂全球文生圖榜單

HiDream-O1-Image-1.5 登頂全球文生圖榜單
PostLinkedIn
⚛️閱讀原文: 量子位
#text-to-image#generative-ai#benchmarkhidream-o1-image-1.5hidream-o1-image-1.5googlenvidia

💡一款新的中國模型在全球文生圖基準測試中超越了 Google 和 Nvidia。

⚡ 30 秒速覽

有什麼變化

HiDream-O1-Image-1.5 在圖像生成領域排名中國第一、全球第二。

為什麼重要

此突破顯示中國生成式模型在全球舞台上的競爭力提升,挑戰了西方科技巨頭在高品質圖像合成領域的主導地位。

下一步行動

評估 HiDream-O1-Image-1.5 與您當前的圖像生成流程,以衡量潛在的品質提升。

誰應關注:Researchers & Academics

關鍵要點

  • HiDream-O1-Image-1.5 在圖像生成領域排名中國第一、全球第二。
  • 該模型在性能上超越了 Google 和 Nvidia 的相關產品。
  • 這標誌著中國 AI 研究在生成式視覺模型領域的重要里程碑。

🧠 深度解析

背景與延伸:來自公開資料,非原文內容。引用 19 個來源。

🔑 增強重點摘要

  • HiDream.ai, the company behind the model, was founded in Beijing, China, in 2023 by Tao Mei, who is a foreign academician of the Canadian Academy of Engineering and former vice president of Jingdong Group.
  • The HiDream-O1-Image-1.5 is a commercial version, while an open-source 8-billion-parameter model, HiDream-O1-Image-Dev, is also available under an MIT license and has shown performance comparable to or surpassing larger open-source models.
  • The model employs a novel Pixel-level Unified Transformer (UiT) architecture that operates directly on raw pixels, eliminating the need for a separate Variational Autoencoder (VAE) or disjoint text encoders, and natively supports high-resolution output up to 2048x2048.
  • HiDream.ai recently secured a Series C funding round in May 2026, with investments from Shenzhen Capital Real Estate Fund, GP Capital, Caixin Capital, and Zhejiang Fuju Investment Management Co., Ltd., indicating strong investor confidence.
  • Beyond text-to-image generation, HiDream-O1-Image-1.5 offers multimodal capabilities including instruction-based image editing, subject-driven personalization, accurate long-text rendering, and storyboard generation within a single architecture.
📊 競品分析▸ Show
ModelDeveloperArchitectureParametersKey FeaturesBenchmarks (ELO Score - Source)Pricing (per 1k images)
HiDream-O1-Image-1.5 (Commercial)HiDream.aiPixel-level Unified Transformer (UiT)200B+ (Pro version)Text-to-image, editing, personalization, long-text rendering, storyboard, native 2K resolution, Reasoning-Driven Prompt AgentRanks 1st in China, 2nd globally (量子位)N/A (Commercial, likely enterprise)
HiDream-O1-Image-Dev (Open-source)HiDream.aiPixel-level Unified Transformer (UiT)8BText-to-image, editing, personalization, long-text rendering, storyboard, native 2K resolution, Reasoning-Driven Prompt Agent1192 ELO (Artificial Analysis)N/A (Open-source, MIT license)
GPT Image 2 (high)OpenAIProprietaryN/AHigh fidelity, best pose/anatomical accuracy, best for mockups1340 ELO (Artificial Analysis)$211.0
Nano Banana 2 (Gemini 3.1 Flash Image Preview)GoogleProprietaryN/AImage generation, editing, multi-turn editing, 14 reference images, 4K resolution, Google Search grounding1260 ELO (Artificial Analysis)$67.0
Cosmos3-Super-Text2Image (agentic)NVIDIAMixture-of-Transformers (Omnimodel)64BPhysical AI, vision reasoning, multimodal generation (text, image, video, sound, action), synthetic data generation1239 ELO (Artificial Analysis)Coming soon
Ideogram 4.0IdeogramProprietary (Open-weight)N/ATop-ranked open-weight, accurate text rendering, LoRA fine-tuningTop spot for open-weight (Artificial Analysis)N/A (API available)

🛠️ 技術深入

  • HiDream-O1-Image is built on a Pixel-level Unified Transformer (UiT) architecture.
  • It operates directly on raw pixels, eliminating the need for an external Variational Autoencoder (VAE) or disjoint text encoders.
  • The model uses a single shared token space for raw pixels, text prompts, and task-specific conditions.
  • The open-source version, HiDream-O1-Image (and its Dev variant), has 8 billion parameters.
  • The commercial 'Pro' version (HiDream-O1-Image-Pro) features over 200 billion parameters.
  • It supports native high-resolution image synthesis up to 2048x2048 pixels.
  • The model incorporates a 'Reasoning-Driven Prompt Agent' that resolves implicit knowledge, layout, and text rendering before generation.
  • It is designed for multiple tasks, including text-to-image generation, instruction-based image editing, subject-driven personalization, long-text rendering, and storyboard generation.
  • The distilled Dev variant can converge in 28 inference steps with a guidance scale of 0, while the full model uses 50 steps with a CFG of 5.0.

🔮 前景展望基於引用來源的 AI 分析

The novel pixel-level architecture could set a new industry standard.
By eliminating VAEs and operating directly on raw pixels, HiDream-O1-Image's architecture addresses common challenges in fine detail and text rendering, potentially influencing future model designs across the industry.
HiDream.ai is poised to strengthen China's position in the global generative AI landscape.
The model's top-tier performance on global benchmarks, combined with significant Series C funding, indicates HiDream.ai's growing influence and competitive edge in advanced AI research and commercialization.
The multimodal capabilities suggest a broader strategic vision beyond static image generation.
With support for image, video, and 3D content generation, HiDream.ai is positioning itself as a comprehensive multimodal content generation and application developer, aiming for a more complete 'world modeling' ability.

時間線

2023
HiDream.ai founded by Tao Mei in Beijing, China
2025-03
Original HiDream-I1, a 17B sparse-MoE DiT model, was released
2026-05-08
HiDream-O1-Image (8B parameters) and its distilled Dev variant, along with the Reasoning-Driven Prompt Agent, were open-sourced under an MIT license
2026-05-10
The technical report for HiDream-O1-Image was made available
2026-05-19
HiDream.ai announced the completion of a Series C funding round
2026-06-10
HiDream-O1-Image-1.5 was reported to rank first in China and second globally on text-to-image benchmarks, surpassing Google and Nvidia
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 量子位

這是摘要,不是原文。去看原站,或訂閱每週簡報。

每週電子報

每週一封,可隨時退訂。