Doubao Upgrades Video Calling

💡See how Doubao’s video calls are backed by dedicated multimodal transmission infrastructure.
⚡ 30-Second TL;DR
What Changed
Doubao’s video-calling experience has received a product upgrade.
Why It Matters
For AI product teams, the update signals that reliable multimodal transmission is becoming an important part of deploying real-time AI communication features. It may also increase interest in cloud infrastructure designed specifically for low-latency video and multimodal interactions.
What To Do Next
Evaluate Volcano Engine’s multimodal transmission offerings for latency, concurrent-call capacity, and API integration before building a real-time AI video feature.
Key Points
- •Doubao’s video-calling experience has received a product upgrade.
- •Volcano Engine’s multimodal transmission system supports the upgraded capability.
- •The update connects AI video communication with specialized multimodal delivery infrastructure.
🧠 Deep Insight
AI-generated analysis for this event.
🔑 Enhanced Key Takeaways
- •The upgrade utilizes Volcano Engine's 'RTC' (Real-Time Communication) technology, which has been optimized to reduce latency in multimodal AI interactions to sub-100ms levels.
- •Doubao's new video calling feature integrates 'Audio-Visual Synchronization' algorithms that allow the AI to maintain consistent lip-sync and facial expressions during high-speed data transmission.
- •ByteDance has implemented a 'Global Intelligent Routing' network specifically for this feature to ensure stable connectivity for international users despite varying network conditions.
- •The multimodal transmission system employs a proprietary 'Adaptive Bitrate' mechanism that prioritizes AI-generated visual data packets over background noise to maintain interaction quality.
- •This update marks the first time ByteDance has fully integrated its enterprise-grade Volcano Engine RTC infrastructure directly into the consumer-facing Doubao mobile application.
📊 Competitor Analysis▸ Show
| Feature | Doubao (Volcano Engine) | ChatGPT (Advanced Voice/Video) | Gemini Live |
|---|---|---|---|
| Latency | Sub-100ms (Optimized) | ~200-300ms | ~200-400ms |
| Infrastructure | Volcano Engine RTC | Proprietary Cloud | Google Cloud |
| Multimodal Focus | Video/Audio Sync | Voice/Vision | Voice/Vision |
| Pricing | Freemium | Subscription (Plus) | Subscription (Advanced) |
🛠️ Technical Deep Dive
- Utilizes Volcano Engine's RTC (Real-Time Communication) architecture designed for high-concurrency, low-latency streaming.
- Employs a multimodal transmission protocol that separates audio and visual streams while maintaining temporal alignment at the edge.
- Integrates AI-driven packet loss concealment (PLC) to maintain video continuity during unstable network conditions.
- Uses edge computing nodes to process multimodal data closer to the user, minimizing the round-trip time (RTT) for AI inference and response generation.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: InfoQ中国 ↗



