🔥較早收集於 8m

有道全面開源 Ziyue 4 多模態與 TTS 引擎

有道全面開源 Ziyue 4 多模態與 TTS 引擎
PostLinkedIn
🔥閱讀原文: 36氪

💡獲取全新的開源多模態與 TTS 引擎,以增強您的 AI 代理感知能力。

⚡ 30-Second TL;DR

有什麼變化

Ziyue 4 支援文本、圖片、音訊的融合交互。

為什麼重要

此舉為開發者提供了多模態與 TTS 任務的開源新選擇,有望降低構建整合式 AI 代理的成本。

下一步行動

從官方倉庫下載 Ziyue 4 開源模型,並將其 TTS 效能與當前業界標準進行基準測試。

誰應關注:Developers & AI Engineers

關鍵要點

  • Ziyue 4 支援文本、圖片、音訊的融合交互。
  • 核心多模態模型正式開源。
  • 語音合成 (TTS) 模型正式開源。
  • 標誌著有道在開源策略上的重大轉向。

🧠 深度解析

Web-grounded analysis with 10 cited sources.

🔑 增強重點摘要

  • Youdao's Ziyue LLM was initially launched in July 2023 as China's first educational large language model, demonstrating its foundational focus on the education sector.
  • Prior to Ziyue 4, Youdao had already established an open-source strategy by releasing models like the Ziyue 3 mathematical model (August 2025), which is China's first open-source inference model for mathematics education, and Ziyue-o1 (January 2025), a lightweight 14-billion-parameter reasoning model for consumer GPUs.
  • The Ziyue LLM powers a range of Youdao's educational products, including the Youdao Dictionary, Youdao Translation, and smart hardware such as the Youdao Dictionary Pen X7 series, indicating a broad application across its ecosystem.
  • Youdao's open-source contributions also include its proprietary RAG engine, QAnything (open-sourced earlier in 2024), and the EmotiVoice TTS engine, which supports multi-voice and prompt-controlled synthesis.
  • This open-source shift aligns with Youdao's broader "AI-native" strategy, which has led to four consecutive quarters of profitability by Q2 2025, driven by commercial success in education and advertising.

🛠️ 技術深入

  • Ziyue-o1 Reasoning Model: This model, open-sourced in January 2025, is a lightweight single model with 14 billion parameters. It is specifically designed for consumer-grade graphics cards, capable of running stably on devices with low video memory. It employs chain-of-thought techniques to provide detailed problem-solving processes and logical reasoning, making its operational thinking closer to human thought processes through "self-talk" and self-correction.
  • Ziyue Translation Large Model 2.0: Upgraded in March 2025, this includes a 14B small-parameter domain-specific model. It utilizes techniques such as large model distillation, model fusion, and Online DPO (Direct Preference Optimization) to enhance translation performance, operational efficiency, accuracy, and fluency while avoiding catastrophic forgetting issues. The model was trained with high-quality translation corpus data meticulously annotated by specialists.
  • Multimodal Capabilities: While specific architectural details for Ziyue 4's multimodal fusion are not detailed, Youdao has open-sourced its EmotiVoice TTS engine, which is a multi-voice and prompt-controlled text-to-speech system. The Ziyue 4 is stated to support integrated interaction across text, image, and audio.

🔮 前景展望AI analysis grounded in cited sources

Youdao will likely see increased adoption of its AI models in the developer community.
Open-sourcing core multimodal and TTS engines lowers the barrier to entry for developers, encouraging integration and innovation within Youdao's ecosystem.
Youdao's position in the EdTech AI market will be further strengthened.
By open-sourcing advanced educational and multimodal AI models, Youdao can foster a community around its technology, potentially leading to more robust and diverse educational applications.
The open-sourcing could accelerate the development of more sophisticated AI-powered learning tools.
Providing full-modal capabilities (text, image, audio) in an open-source format allows other developers and institutions to build more comprehensive and interactive learning experiences.

時間線

2006
NetEase Youdao established as a subsidiary of NetEase Inc.
2023-07
Youdao launched "Ziyue," China's first educational LLM
2024-Early
Youdao open-sourced its proprietary RAG engine, QAnything
2025-01
Youdao open-sourced "ZiYue-o1," a 14B parameter reasoning model for consumer GPUs
2025-08
Youdao open-sourced the "ZiYue 3" mathematical model, China's first open-source inference model for math education
2026-05
Youdao Open-Sources Ziyue 4 Multimodal and TTS Engines

📎 來源 (10)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. prnewswire.com
  2. yicaiglobal.com
  3. reddit.com
  4. moomoo.com
  5. tmtpost.com
  6. aibase.com
  7. tmtpost.com
  8. github.com
  9. startupintros.com
  10. wikipedia.org
📰

AI 週報

閱讀本週精選 AI 大事摘要 →

👉相關動態

AI 策展新聞聚合。所有內容版權歸原始發布者所有。
原始來源: 36氪