🔥Stalecollected in 8m

Youdao Open-Sources Ziyue 4 Multimodal and TTS Engines

Youdao Open-Sources Ziyue 4 Multimodal and TTS Engines
PostLinkedIn
🔥Read original on 36氪

💡Access new open-source multimodal and TTS engines to enhance your AI agent's sensory capabilities.

⚡ 30-Second TL;DR

What Changed

Ziyue 4 supports integrated interaction across text, image, and audio.

Why It Matters

This release provides developers with new open-source alternatives for multimodal and TTS tasks, potentially lowering costs for building integrated AI agents.

What To Do Next

Download the Ziyue 4 open-source models from their repository to benchmark their TTS performance against current industry standards.

Who should care:Developers & AI Engineers

Key Points

  • Ziyue 4 supports integrated interaction across text, image, and audio.
  • Core multimodal model is now open-source.
  • TTS (Text-to-Speech) engine is now open-source.
  • Represents a significant shift toward open-source strategy for Youdao.

🧠 Deep Insight

Web-grounded analysis with 10 cited sources.

🔑 Enhanced Key Takeaways

  • Youdao's Ziyue LLM was initially launched in July 2023 as China's first educational large language model, demonstrating its foundational focus on the education sector.
  • Prior to Ziyue 4, Youdao had already established an open-source strategy by releasing models like the Ziyue 3 mathematical model (August 2025), which is China's first open-source inference model for mathematics education, and Ziyue-o1 (January 2025), a lightweight 14-billion-parameter reasoning model for consumer GPUs.
  • The Ziyue LLM powers a range of Youdao's educational products, including the Youdao Dictionary, Youdao Translation, and smart hardware such as the Youdao Dictionary Pen X7 series, indicating a broad application across its ecosystem.
  • Youdao's open-source contributions also include its proprietary RAG engine, QAnything (open-sourced earlier in 2024), and the EmotiVoice TTS engine, which supports multi-voice and prompt-controlled synthesis.
  • This open-source shift aligns with Youdao's broader "AI-native" strategy, which has led to four consecutive quarters of profitability by Q2 2025, driven by commercial success in education and advertising.

🛠️ Technical Deep Dive

  • Ziyue-o1 Reasoning Model: This model, open-sourced in January 2025, is a lightweight single model with 14 billion parameters. It is specifically designed for consumer-grade graphics cards, capable of running stably on devices with low video memory. It employs chain-of-thought techniques to provide detailed problem-solving processes and logical reasoning, making its operational thinking closer to human thought processes through "self-talk" and self-correction.
  • Ziyue Translation Large Model 2.0: Upgraded in March 2025, this includes a 14B small-parameter domain-specific model. It utilizes techniques such as large model distillation, model fusion, and Online DPO (Direct Preference Optimization) to enhance translation performance, operational efficiency, accuracy, and fluency while avoiding catastrophic forgetting issues. The model was trained with high-quality translation corpus data meticulously annotated by specialists.
  • Multimodal Capabilities: While specific architectural details for Ziyue 4's multimodal fusion are not detailed, Youdao has open-sourced its EmotiVoice TTS engine, which is a multi-voice and prompt-controlled text-to-speech system. The Ziyue 4 is stated to support integrated interaction across text, image, and audio.

🔮 Future ImplicationsAI analysis grounded in cited sources

Youdao will likely see increased adoption of its AI models in the developer community.
Open-sourcing core multimodal and TTS engines lowers the barrier to entry for developers, encouraging integration and innovation within Youdao's ecosystem.
Youdao's position in the EdTech AI market will be further strengthened.
By open-sourcing advanced educational and multimodal AI models, Youdao can foster a community around its technology, potentially leading to more robust and diverse educational applications.
The open-sourcing could accelerate the development of more sophisticated AI-powered learning tools.
Providing full-modal capabilities (text, image, audio) in an open-source format allows other developers and institutions to build more comprehensive and interactive learning experiences.

Timeline

2006
NetEase Youdao established as a subsidiary of NetEase Inc.
2023-07
Youdao launched "Ziyue," China's first educational LLM
2024-Early
Youdao open-sourced its proprietary RAG engine, QAnything
2025-01
Youdao open-sourced "ZiYue-o1," a 14B parameter reasoning model for consumer GPUs
2025-08
Youdao open-sourced the "ZiYue 3" mathematical model, China's first open-source inference model for math education
2026-05
Youdao Open-Sources Ziyue 4 Multimodal and TTS Engines

📎 Sources (10)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. prnewswire.com
  2. yicaiglobal.com
  3. reddit.com
  4. moomoo.com
  5. tmtpost.com
  6. aibase.com
  7. tmtpost.com
  8. github.com
  9. startupintros.com
  10. wikipedia.org
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: 36氪