New Chinese Open-Source Framework Enables 5-Minute AI Videos

💡First open-source framework to enable stable 5-minute AI video generation with real-time super-resolution.
⚡ 30-Second TL;DR
What Changed
Supports stable generation of 5-minute long AI videos
Why It Matters
This development challenges the dominance of international video models like Sora and Kling by providing a high-performance open-source alternative for developers.
What To Do Next
Visit the project's GitHub repository to benchmark its temporal consistency against existing models like Stable Video Diffusion.
Key Points
- •Supports stable generation of 5-minute long AI videos
- •Features high temporal consistency and low latency
- •Includes integrated real-time super-resolution capabilities
🧠 Deep Insight
Web-grounded analysis with 10 cited sources.
🔑 Enhanced Key Takeaways
- •The new open-source AI framework is officially named JoyAI-Echo and has been developed by JD.com, a major Chinese tech company.
- •JoyAI-Echo is distinguished as the first open-source model capable of generating multi-shot videos up to five minutes long while maintaining 100% consistency in both character appearance and vocal timbre.
- •The framework incorporates a 'director agent' feature, allowing users to generate entire screenplays, scene prompts, and shot lists through conversational input or structured JSON prompts.
- •It achieves a significant 7.5x inference speedup for streaming generation through its proprietary Distribution Matching Distillation (DMD) technology.
- •For optimal performance, JoyAI-Echo requires substantial hardware, with recommendations for enterprise-grade GPUs like the H100 or A100 with 48 to 80 GB of VRAM, though it can run on consumer-grade cards like an RTX 5090 (32GB VRAM) at lower resolutions or shorter clip lengths.
📊 Competitor Analysis▸ Show
| Feature / Model | JoyAI-Echo (JD.com) | Open-Sora 2.5 | Wan 2.2 (Alibaba) | HunyuanVideo (Tencent) | LTX-2.3 | Sora (OpenAI) |
|---|---|---|---|---|---|---|
| Type | Open-source | Open-source | Open-source | Open-source | Open-source | Closed-source |
| Max Video Length | 5 minutes | 2 minutes | Up to 81 frames (Wan 2.1), 5-8 seconds | ~5 seconds (129 frames) | Not specified for long-form | "Several minutes" for short clips, struggles with long-form consistency |
| Temporal Consistency | High (character appearance & vocal timbre) | Consistent character physics | Better over 5-8 sec clips | High generation stability | Addresses consistency via 3D RoPE, STG, temporal upsampling | Struggles with consistency across long-form |
| Real-time Super-resolution | Integrated (1K resolution on the fly) | Not specified | Not specified | Separate SR step (1080p) | Not specified | Not specified |
| Latency/Speed | 7.5x inference speedup via DMD | Not specified | Not specified | Generation speed in minutes for 720p/129 frames | Distilled models like LTX-2.3 are faster | "Several minutes" for short clips |
| Hardware Requirements | Enterprise H100/A100 (48-80GB VRAM) recommended; RTX 5090 (32GB VRAM) for lower settings | Not specified | 16GB VRAM for 14B model | 60GB+ VRAM for full precision; 24GB for quantized | Not specified | Not specified |
| Unique Features | Director agent for screenplay/shot lists; simultaneous audio/video output | Excels at complex spatial-temporal data | Good motion quality, realistic camera movements | Unified multitask DiT (T2I, T2V, I2V); Sparse Attention | 3D RoPE positional encoding, Spatio-Temporal Guidance (STG) | High quality, long-form possible but inaccessible |
🛠️ Technical Deep Dive
- Model Name: JoyAI-Echo
- Developer: JD.com
- Core Architecture: Employs a "slot-paired cross-modal memory bank" to maintain long-term consistency in both visual appearance and vocal timbre across extended video sequences.
- Temporal Consistency Mechanism: The slot-paired cross-modal memory bank specifically addresses identity drift by locking together slots, preventing context forgetting over long horizons.
- Inference Acceleration: Utilizes Distribution Matching Distillation (DMD) technology, which acts as an algorithmic shortcut, resulting in a 7.5x faster inference speed for streaming generation.
- Super-Resolution: Integrates a lightweight super-resolution module capable of upscaling native generation to 1K resolution in real-time without introducing streaming lag.
- Multimodal Output: Capable of pumping out audio and video simultaneously from a single pipeline.
- Input Flexibility: Supports conversational input for its director agent to generate screenplays and shot lists, or can accept structured JSON prompts.
- Model Size: The raw model weights (safe tensors) are approximately 46.1 GB.
- Hardware Requirements: Generating a 10-second video at 736p and 25 frames per second can peak at 46-50 GB of VRAM. Recommended hardware includes a single Enterprise H100 or A100 GPU with 48 or 80 GB of RAM. Consumer-grade GPUs like an RTX 5090 (32 GB VRAM) can run the model by scaling down resolution to 480p, reducing frame rate, or shortening clip length.
🔮 Future ImplicationsAI analysis grounded in cited sources
⏳ Timeline
📎 Sources (10)
Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.
Weekly AI Recap
Read this week's curated digest of top AI events →
👉Related Updates
AI-curated news aggregator. All content rights belong to original publishers.
Original source: 量子位 ↗