較早收集於 7h

智象未來發布超兩千億參數原生全模態模型 HiDream-O1-Image-Pro

智象未來發布超兩千億參數原生全模態模型 HiDream-O1-Image-Pro
PostLinkedIn
閱讀原文: 雷峰网

💡一款超過 2000 億參數的原生全模態模型,透過統一架構挑戰現有 SOTA 基準測試。

⚡ 30-Second TL;DR

有什麼變化

發布基於 UiT 架構、超過 2000 億參數的閉源模型 HiDream-O1-Image-Pro。

為什麼重要

向「原生全模態」架構的轉變,顯示業界正從模組化拼接模型轉向具備更強物理推理與空間理解能力的統一世界模型。

下一步行動

評估 HiDream-O1-Image-Pro 在高保真文生圖工作流中的表現,特別是針對需要精準文字渲染與複雜多主體合成的場景。

誰應關注:Developers & AI Engineers

關鍵要點

  • 發布基於 UiT 架構、超過 2000 億參數的閉源模型 HiDream-O1-Image-Pro。
  • 在複雜文字渲染、多主體個性化及指令編輯任務中達到 SOTA 水準。
  • 獲得包括深創投、金浦投資等多家機構的新一輪融資。
  • 採用「1+1+3」業務架構,涵蓋底層模型、能力中台及商業營銷、影視創作、社媒創作三大應用場景。

🧠 深度解析

Web-grounded analysis with 18 cited sources.

🔑 增強重點摘要

  • HiDream.ai was founded in March 2023 by Dr. Mei Tao, a former JD.com VP and 2025 ACM Fellow, and is headquartered in Beijing, China.
  • The company has secured over 500 million CNY (approximately $70M USD) in funding, with recent investors including Shenzhen Capital Group (SCGC), GP Capital (Jinpu Investment), Caixin Capital, Fuju Investment, Oriental Fortune Capital, Anhui Industrial Investment, and Fenghua Capital.
  • Prior to the 200B+ Pro model, HiDream.ai released an 8-billion-parameter open-source text-to-image model, HiDream-O1-Image (also referred to as HiDream-I1), which demonstrated superior performance over larger models like FLUX.2 [dev] (32B parameters) across multiple image quality benchmarks.
  • The Unified Transformer (UiT) architecture employed by HiDream-O1-Image-Pro represents a 'native full-modal' approach, aiming to overcome the limitations of fragmented architectures that rely on separate text encoders and external VAEs by integrating various modalities into a shared latent space.
  • HiDream.ai's product ecosystem extends beyond image generation to include HiDream-E1 for image editing, vivago.ai as a consumer AIGC platform, and PixMaker for marketing content, with plans to release a minute-long video generation model capable of synchronized audio and visual output.
📊 競品分析▸ Show
Feature/BenchmarkHiDream-O1-Image (8B)FLUX.2 [dev] (32B)DALL-E 3Ideogram v3
Parameter Count8 Billion32 BillionProprietaryProprietary
ArchitecturePixel-level Unified Transformer (UiT), no VAEDiffusion Transformer (implied VAE)ProprietaryProprietary
LicenseMIT (Open Weights)Developer VersionProprietaryProprietary
GenEval (overall)0.900.87--
DPG-Bench (overall)89.8387.5783.50-
HPSv3 (overall)10.379.28--
CVTG-2K (average)0.91280.8926--
LongText-EN0.9790.963--
LongText-ZH0.9780.757--
Text RenderingSOTA performance--Strong on text rendering and typography
Image GenerationSOTA performance---
Multi-task EditingSOTA performance---

Note: The HiDream-O1-Image-Pro (200B+) is a closed-source model that has reportedly surpassed mainstream open-source models like Z-Image Turbo, Qwen-Image, and FLUX.2 [dev] in benchmarks, but specific comparative scores for the Pro version are not publicly detailed.

🛠️ 技術深入

  • Unified Transformer (UiT) Architecture: HiDream-O1-Image-Pro is built on a new generation native full-modal UiT architecture, which aims to provide a unified modeling framework for images, videos, text, audio, and potentially action and embodied data.
  • Pixel-space Diffusion Transformer: The architecture pioneers a paradigm shift from modular designs (relying on disjoint text encoders and external VAEs) to an end-to-end in-context visual generation engine, specifically using a pixel-space Diffusion Transformer that removes the VAE bottleneck.
  • Scalability: The architecture has been successfully scaled up to over 200 billion parameters for the HiDream-O1-Image-Pro, demonstrating immense scalability.
  • HiDream-I1 (8B/17B) Architectural Insights (predecessor/related model):
    • Employs a Mixture of Experts (MoE) architecture within its Diffusion Transformer (DiT) backbone, combining dual-flow MMDiT blocks with single-flow DiT blocks for efficient resource allocation via dynamic routing.
    • Integrates multiple text encoders, including OpenCLIP ViT-bigG, OpenAI CLIP ViT-L, T5-XXL, and Llama-3.1-8B-Instruct, to significantly enhance semantic understanding.
    • Features a dual-stream then single-stream design for cross-attention: image latent tokens and text tokens are processed in parallel before merging into a single stream where a unified transformer layer attends over the concatenated sequence.
    • Conditioning, such as pooled CLIP embeddings and diffusion timesteps, is injected via adaptive layer normalization (AdaLN) in each transformer block.
    • The 8B model supports native generation up to 2048x2048 pixels.

🔮 前景展望AI analysis grounded in cited sources

HiDream.ai's 'native full-modal' UiT architecture could become a foundational approach for next-generation multimodal AI.
By unifying various modalities into a single framework and addressing the limitations of fragmented architectures, it aims to enable more comprehensive 'world modeling' capabilities, potentially influencing future AI development.
The company's strategic focus on enterprise-focused intelligent agents and global market expansion indicates a strong push for broad commercial adoption beyond creative content generation.
Recent funding rounds are specifically allocated for the research and development of next-generation native multimodal world models, enterprise service intelligent agents, and global market expansion, signaling a move towards diverse industrial applications.

時間線

2023-03
HiDream.ai founded by Dr. Mei Tao.
2023-12
HiDream.ai receives funding led by iFLYTEK CO.,LTD.
2024-05
HiDream.ai launches the world's first publicly available DiT-architecture video generation model.
2025-04
HiDream.ai officially open-sources HiDream-I1, a 17-billion-parameter text-to-image model.
2026-04
HiDream.ai closes a new funding round exceeding 500 million yuan (approx. $70M USD).
2026-05
HiDream.ai releases the 8-billion-parameter open-source HiDream-O1-Image model and its technical report, and officially releases the 200B+ parameter closed-source HiDream-O1-Image-Pro model, announcing another round of hundred-million-level financing.
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 雷峰网