PulseAugur
EN
LIVE 15:30:10

Meituan releases LongCat-Video-Avatar 1.5 for stable, audio-driven video generation

Meituan has released LongCat-Video-Avatar 1.5, an open-source framework for generating audio-driven human videos. This upgraded version emphasizes production-readiness and stability, featuring an improved Whisper-Large audio encoder for more natural lip-syncing and robust long-video generation with consistent identity. The model supports various tasks like Audio-Text-to-Video and Video Continuation, generalizing across diverse styles and conditions, and achieves efficient 8-step inference. AI

IMPACT Accelerates development of audio-driven video generation tools with a focus on stability and efficiency.

RANK_REASON New open-source model release from a significant AI lab (Meituan).

Read on Hugging Face Trending Models →

AI-generated summary · Google Gemini · from 5 sources. How we write summaries →

Meituan releases LongCat-Video-Avatar 1.5 for stable, audio-driven video generation

COVERAGE [5]

  1. Hugging Face Trending Models TIER_1 (CA) · meituan-longcat ·

    meituan-longcat/LongCat-Video-Avatar-1.5

    0 downloads · 132 likes

  2. arXiv cs.CV TIER_1 English(EN) · Meituan LongCat Team, Xunliang Cai, Meng Cheng, Feng Gao, Zhe Kong, Jiamu Li, Le Li, Weiheng Li, Hongyu Liu, Shuai Tan, Xiaoming Wei, Tianyu Yang, Yong Zhang ·

    LongCat-Video-Avatar 1.5 Technical Report

    arXiv:2605.26486v1 Announce Type: new Abstract: Despite advances in audio-driven video generation, achieving commercial-grade stability remains challenging. We present LongCat-Video-Avatar 1.5, an upgraded open-source framework prioritizing systematic engineering and production-r…

  3. Pandaily TIER_1 English(EN) · [email protected] (Pandaily) ·

    Meituan Open-Sources LongCat-Video-Avatar 1.5: Photorealistic Digital Human Video Framework

    Meituan releases version 1.5 of its open-source digital human video generation framework, achieving state-of-the-art lip-sync accuracy with just 8 inference steps.

  4. Mastodon — fosstodon.org TIER_1 Italiano(IT) · [email protected] ·

    🧠 LongCat-Video-Avatar 1.5 is a new open-source framework by #Meituan for generating audio-driven video avatars, with a strong focus on stability

    🧠 LongCat-Video-Avatar 1.5 è un nuovo framework open source di # Meituan per generare video avatar guidati dall’audio, con un forte orientamento alla stabilità e all’utilizzo in contesti produttivi. 👉 I dettagli: https://www. linkedin.com/posts/alessiopoma ro_meituan-ai-genai-act…

  5. r/StableDiffusion TIER_2 English(EN) · /u/Turbulent_Corner9895 ·

    LongCat-Video-Avatar 1.5 Release

    <table> <tr><td> <a href="https://www.reddit.com/r/StableDiffusion/comments/1tm5oxh/longcatvideoavatar_15_release/"> <img alt="LongCat-Video-Avatar 1.5 Release" src="https://preview.redd.it/j7ay6s16j13h1.png?width=640&amp;crop=smart&amp;auto=webp&amp;s=e2eac6efeee2e3d8dc34d88c058…