Youdao Open-Sources Ziyue 4 Multimodal and TTS Engines
💡Access new open-source multimodal and TTS engines to enhance your AI agent's sensory capabilities.
⚡ 30-Second TL;DR
What Changed
Ziyue 4 supports integrated interaction across text, image, and audio.
Why It Matters
This release provides developers with new open-source alternatives for multimodal and TTS tasks, potentially lowering costs for building integrated AI agents.
What To Do Next
Download the Ziyue 4 open-source models from their repository to benchmark their TTS performance against current industry standards.
Key Points
- •Ziyue 4 supports integrated interaction across text, image, and audio.
- •Core multimodal model is now open-source.
- •TTS (Text-to-Speech) engine is now open-source.
- •Represents a significant shift toward open-source strategy for Youdao.
🧠 Deep Insight
Web-grounded analysis with 10 cited sources.
🔑 Enhanced Key Takeaways
- •Youdao's Ziyue LLM was initially launched in July 2023 as China's first educational large language model, demonstrating its foundational focus on the education sector.
- •Prior to Ziyue 4, Youdao had already established an open-source strategy by releasing models like the Ziyue 3 mathematical model (August 2025), which is China's first open-source inference model for mathematics education, and Ziyue-o1 (January 2025), a lightweight 14-billion-parameter reasoning model for consumer GPUs.
- •The Ziyue LLM powers a range of Youdao's educational products, including the Youdao Dictionary, Youdao Translation, and smart hardware such as the Youdao Dictionary Pen X7 series, indicating a broad application across its ecosystem.
- •Youdao's open-source contributions also include its proprietary RAG engine, QAnything (open-sourced earlier in 2024), and the EmotiVoice TTS engine, which supports multi-voice and prompt-controlled synthesis.
- •This open-source shift aligns with Youdao's broader "AI-native" strategy, which has led to four consecutive quarters of profitability by Q2 2025, driven by commercial success in education and advertising.
🛠️ Technical Deep Dive
- Ziyue-o1 Reasoning Model: This model, open-sourced in January 2025, is a lightweight single model with 14 billion parameters. It is specifically designed for consumer-grade graphics cards, capable of running stably on devices with low video memory. It employs chain-of-thought techniques to provide detailed problem-solving processes and logical reasoning, making its operational thinking closer to human thought processes through "self-talk" and self-correction.
- Ziyue Translation Large Model 2.0: Upgraded in March 2025, this includes a 14B small-parameter domain-specific model. It utilizes techniques such as large model distillation, model fusion, and Online DPO (Direct Preference Optimization) to enhance translation performance, operational efficiency, accuracy, and fluency while avoiding catastrophic forgetting issues. The model was trained with high-quality translation corpus data meticulously annotated by specialists.
- Multimodal Capabilities: While specific architectural details for Ziyue 4's multimodal fusion are not detailed, Youdao has open-sourced its EmotiVoice TTS engine, which is a multi-voice and prompt-controlled text-to-speech system. The Ziyue 4 is stated to support integrated interaction across text, image, and audio.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (10)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 36氪 ↗