๐ŸผStalecollected in 28m

Meituan Open-Sources LongCat-Video-Avatar 1.5 for Digital Humans

Meituan Open-Sources LongCat-Video-Avatar 1.5 for Digital Humans
PostLinkedIn
๐ŸผRead original on Pandaily

๐Ÿ’กHigh-efficiency, open-source digital human framework with state-of-the-art lip-sync performance.

โšก 30-Second TL;DR

What Changed

Achieves state-of-the-art lip-sync accuracy in digital human generation.

Why It Matters

This release lowers the barrier for developers to create high-quality, real-time digital avatars. It significantly improves the efficiency of lip-syncing in generative video pipelines.

What To Do Next

Clone the LongCat-Video-Avatar repository and benchmark the 8-step inference latency against your current lip-sync solution.

Who should care:Developers & AI Engineers

Key Points

  • โ€ขAchieves state-of-the-art lip-sync accuracy in digital human generation.
  • โ€ขOptimized for efficiency with only 8 inference steps required.
  • โ€ขOpen-source release allows developers to integrate photorealistic avatars.

๐Ÿง  Deep Insight

Web-grounded analysis with 13 cited sources.

๐Ÿ”‘ Enhanced Key Takeaways

  • โ€ขLongCat-Video-Avatar 1.5 integrates the Whisper-Large audio encoder, enabling enhanced lip-sync accuracy and style generalization across 99 languages, a significant upgrade from its predecessor's Wav2Vec2.
  • โ€ขThe framework achieves a substantial efficiency gain, boasting approximately 15 times faster inference speed by reducing generation steps from 50 to 8 through Distribution Matching Distillation (DMD).
  • โ€ขThe model demonstrates robust generalization capabilities, supporting the generation of digital humans in diverse styles, including realistic human portraits, anime characters, virtual idols, and even animals, while also handling multi-person interactions.
  • โ€ขBuilt upon the LongCat-Video foundation, version 1.5 is designed for long-duration video generation, capable of producing coherent content up to 15 minutes, addressing a key limitation in temporal consistency for AI video.
  • โ€ขIt operates with a 'one model for multiple tasks' design, natively supporting Audio-Text-to-Video (AT2V), Audio-Text-Image-to-Video (ATI2V), and video continuation, offering versatile application scenarios.
๐Ÿ“Š Competitor Analysisโ–ธ Show

While specific pricing details are not consistently available for all platforms, especially open-source ones, a comparison of features and benchmarks can be made:

Feature / PlatformMeituan LongCat-Video-Avatar 1.5HeyGenKling AI / Avatar 2.0ByteDance OmniHuman-1Dubly.AI
Lip-Sync AccuracyState-of-the-art, Whisper-Large encoder, 99 languagesStrong, for personalized/translated videosImpressive, strong for dialogueAccurateHighest benchmark (96.4), handles complex footage
Efficiency8 inference steps, ~15x speedup via DMDFast generation timesFast generation timesNot specifiedNot specified
Long-Video StabilityUp to 15 minutes, temporal consistency, identity preservationUp to 3 minutesImproves across long-form videosNot specifiedNot specified
Subject GeneralizationReal humans, anime, virtual idols, animals, multi-personTalking digital twinsPhotorealistic humansPortraits, full body, cartoons, stylized avatarsReal video footage
LicensingOpen-source (MIT License for academic use, commercial licensing available)Commercial platformCommercial platformResearch modelCommercial platform
BenchmarksLeads EvalTalker evaluations for anthropomorphism76.8 (independent benchmark)Best-in-class for realistic human generationNot specified96.4 (independent benchmark)
Unique FeaturesMulti-segment rolling inference, GRPO alignment, Cross-Chunk Latent StitchingPersonalized & translated videos, overview outline before generationBest-in-class for realistic human faces/movementsSingle image + motion signalsGDPR-compliant servers, data not used for AI training

๐Ÿ› ๏ธ Technical Deep Dive

  • Base Model: Built on the LongCat-Video Diffusion Transformer (DiT) architecture, which serves as a foundational video generation model.
  • Audio Encoder: Upgraded from Wav2Vec2 to Whisper-Large (specifically Whisper-large-v3) for superior audio feature extraction, leading to more precise lip synchronization and multilingual support.
  • Inference Efficiency: Utilizes Distribution Matching Distillation (DMD) to compress the generation process from 50 steps to just 8, resulting in an approximate 15x speedup.
  • Memory Optimization: Employs a shared base model combined with multiple LoRA (Low-Rank Adaptation) adapters to significantly reduce VRAM usage compared to traditional three-model parallel deployment schemes. INT8 quantization is also supported for further VRAM reduction.
  • Long Video Stability: Incorporates a hierarchical attention mechanism (fine-grained, medium-scale, coarse attention), Cross-Chunk Latent Stitching, and Reference Skip Attention to maintain visual consistency, narrative coherence, and identity over extended video durations, capable of generating up to 15 minutes of coherent content.
  • Quality and Alignment: Leverages Group Relative Policy Optimization (GRPO) with frame-level human preference rewards to ensure accurate coordination of speech, lip motion, facial expressions, head pose, and body movements.
  • Silent Segment Naturalness: Features Disentangled Unconditional Guidance to generate natural micro-expressions and body dynamics during silent periods, preventing unnatural stillness.
  • Supported Tasks: Natively supports Audio-Text-to-Video (AT2V), Audio-Text-Image-to-Video (ATI2V), and video continuation, accommodating both single-stream and multi-stream audio inputs.
  • Data Pipeline: Utilizes a high-quality, multi-stage data processing workflow involving offline annotation (extracting face keypoints, person count, body composition, audio-visual sync) and online validation (filtering low-quality segments).

๐Ÿ”ฎ Future ImplicationsAI analysis grounded in cited sources

Increased adoption of AI-generated content in commercial production.
The model's 'commercial-grade' quality, enhanced efficiency, and open-source nature with commercial licensing significantly lower the barriers for businesses to integrate advanced digital human technology into their production workflows.
Acceleration of multimodal AI agent development.
Meituan's broader LongCat family strategy aims to build intelligent agents by combining vision (LongCat-Video) and reasoning (LongCat-Flash), suggesting future integration for more sophisticated AI applications that can both understand and generate complex multimodal content.
Democratization of advanced digital human creation.
The open-source release of a state-of-the-art framework allows a wider range of developers, researchers, and creators to access, modify, and build upon cutting-edge technology, fostering innovation beyond large corporate entities.

โณ Timeline

2010-03
Meituan.com founded by Wang Xing, focusing on group-buying services.
2015-10
Meituan merges with Dianping, a major competitor in local services.
2020-09
Company officially rebrands from 'Meituan Dianping' to 'Meituan'.
2023
Meituan acquires AI startup Light Year for US$281 million, marking a significant entry into the AI sector.
2025-09
Meituan releases its first open-source Large Language Model (LLM), LongCat-Flash-Chat.
2025-10
Meituan releases LongCat-Video, a foundational video generation model.
2025-12
Meituan releases LongCat-Video-Avatar (v1.0), an audio-driven character animation model.
2026-05
Meituan open-sources LongCat-Video-Avatar 1.5, an upgraded framework for audio-driven human video generation.

๐Ÿ“Ž Sources (13)

Factual claims are grounded in the sources below. Forward-looking analysis is AI-generated interpretation.

  1. aifilms.ai
  2. phemex.com
  3. longcatai.org
  4. knightli.com
  5. github.com
  6. reddit.com
  7. youtube.com
  8. longcatai.net
  9. wavespeed.ai
  10. runware.ai
  11. manus.im
  12. aijourn.com
  13. crepal.ai
๐Ÿ“ฐ

Weekly AI Recap

Read this week's curated digest of top AI events โ†’

๐Ÿ‘‰Related Updates

AI-curated news aggregator. All content rights belong to original publishers.
Original source: Pandaily โ†—