⚛️Stalecollected in 78m

New Chinese Open-Source Framework Enables 5-Minute AI Videos

New Chinese Open-Source Framework Enables 5-Minute AI Videos
PostLinkedIn
⚛️Read original on 量子位
#video-gen#open-source#long-form-video国产开源视频生成框架github

💡First open-source framework to enable stable 5-minute AI video generation with real-time super-resolution.

⚡ 30-Second TL;DR

What Changed

Supports stable generation of 5-minute long AI videos

Why It Matters

This development challenges the dominance of international video models like Sora and Kling by providing a high-performance open-source alternative for developers.

What To Do Next

Visit the project's GitHub repository to benchmark its temporal consistency against existing models like Stable Video Diffusion.

Who should care:Developers & AI Engineers

Key Points

  • Supports stable generation of 5-minute long AI videos
  • Features high temporal consistency and low latency
  • Includes integrated real-time super-resolution capabilities

🧠 Deep Insight

Web-grounded analysis with 10 cited sources.

🔑 Enhanced Key Takeaways

  • The new open-source AI framework is officially named JoyAI-Echo and has been developed by JD.com, a major Chinese tech company.
  • JoyAI-Echo is distinguished as the first open-source model capable of generating multi-shot videos up to five minutes long while maintaining 100% consistency in both character appearance and vocal timbre.
  • The framework incorporates a 'director agent' feature, allowing users to generate entire screenplays, scene prompts, and shot lists through conversational input or structured JSON prompts.
  • It achieves a significant 7.5x inference speedup for streaming generation through its proprietary Distribution Matching Distillation (DMD) technology.
  • For optimal performance, JoyAI-Echo requires substantial hardware, with recommendations for enterprise-grade GPUs like the H100 or A100 with 48 to 80 GB of VRAM, though it can run on consumer-grade cards like an RTX 5090 (32GB VRAM) at lower resolutions or shorter clip lengths.
📊 Competitor Analysis▸ Show
Feature / ModelJoyAI-Echo (JD.com)Open-Sora 2.5Wan 2.2 (Alibaba)HunyuanVideo (Tencent)LTX-2.3Sora (OpenAI)
TypeOpen-sourceOpen-sourceOpen-sourceOpen-sourceOpen-sourceClosed-source
Max Video Length5 minutes2 minutesUp to 81 frames (Wan 2.1), 5-8 seconds~5 seconds (129 frames)Not specified for long-form"Several minutes" for short clips, struggles with long-form consistency
Temporal ConsistencyHigh (character appearance & vocal timbre)Consistent character physicsBetter over 5-8 sec clipsHigh generation stabilityAddresses consistency via 3D RoPE, STG, temporal upsamplingStruggles with consistency across long-form
Real-time Super-resolutionIntegrated (1K resolution on the fly)Not specifiedNot specifiedSeparate SR step (1080p)Not specifiedNot specified
Latency/Speed7.5x inference speedup via DMDNot specifiedNot specifiedGeneration speed in minutes for 720p/129 framesDistilled models like LTX-2.3 are faster"Several minutes" for short clips
Hardware RequirementsEnterprise H100/A100 (48-80GB VRAM) recommended; RTX 5090 (32GB VRAM) for lower settingsNot specified16GB VRAM for 14B model60GB+ VRAM for full precision; 24GB for quantizedNot specifiedNot specified
Unique FeaturesDirector agent for screenplay/shot lists; simultaneous audio/video outputExcels at complex spatial-temporal dataGood motion quality, realistic camera movementsUnified multitask DiT (T2I, T2V, I2V); Sparse Attention3D RoPE positional encoding, Spatio-Temporal Guidance (STG)High quality, long-form possible but inaccessible

🛠️ Technical Deep Dive

  • Model Name: JoyAI-Echo
  • Developer: JD.com
  • Core Architecture: Employs a "slot-paired cross-modal memory bank" to maintain long-term consistency in both visual appearance and vocal timbre across extended video sequences.
  • Temporal Consistency Mechanism: The slot-paired cross-modal memory bank specifically addresses identity drift by locking together slots, preventing context forgetting over long horizons.
  • Inference Acceleration: Utilizes Distribution Matching Distillation (DMD) technology, which acts as an algorithmic shortcut, resulting in a 7.5x faster inference speed for streaming generation.
  • Super-Resolution: Integrates a lightweight super-resolution module capable of upscaling native generation to 1K resolution in real-time without introducing streaming lag.
  • Multimodal Output: Capable of pumping out audio and video simultaneously from a single pipeline.
  • Input Flexibility: Supports conversational input for its director agent to generate screenplays and shot lists, or can accept structured JSON prompts.
  • Model Size: The raw model weights (safe tensors) are approximately 46.1 GB.
  • Hardware Requirements: Generating a 10-second video at 736p and 25 frames per second can peak at 46-50 GB of VRAM. Recommended hardware includes a single Enterprise H100 or A100 GPU with 48 or 80 GB of RAM. Consumer-grade GPUs like an RTX 5090 (32 GB VRAM) can run the model by scaling down resolution to 480p, reducing frame rate, or shortening clip length.

🔮 Future ImplicationsAI analysis grounded in cited sources

Democratization of long-form AI video creation will accelerate.
As an open-source framework with advanced capabilities for long-form, high-consistency video generation, JoyAI-Echo lowers the barrier to entry for creators and developers who may not have access to or budget for proprietary solutions.
AI-driven content production workflows will become significantly more efficient.
The combination of a director agent for automated screenplay generation and a 7.5x inference speedup for streaming generation will drastically reduce the time and complexity involved in producing complex video content, from conceptualization to final output.
Multimodal consistency will become a standard expectation for advanced AI video models.
JoyAI-Echo's ability to maintain both visual and vocal timbre consistency across extended video lengths sets a new benchmark, pushing future AI video models to prioritize and integrate similar multimodal coherence capabilities.

Timeline

2026-06
JD.com releases JoyAI-Echo, a new open-source AI framework for 5-minute, high-consistency AI video generation with real-time super-resolution.

📎 Sources (10)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. youtube.com
  2. digen.ai
  3. ltx.io
  4. crepal.ai
  5. ltx.io
  6. scmp.com
  7. colossyan.com
  8. fal.ai
  9. medium.com
  10. modal.com
📰

Weekly AI Recap

Read this week's curated digest of top AI events →

👉Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: 量子位