Meituan has released LongCat-Video-Avatar 1.5, an open-source framework for generating audio-driven human videos. This upgraded version emphasizes production-readiness and stability, featuring an improved Whisper-Large audio encoder for more natural lip-syncing and robust long-video generation with consistent identity. The model supports various tasks like Audio-Text-to-Video and Video Continuation, generalizing across diverse styles and conditions, and achieves efficient 8-step inference. AI
IMPACT Accelerates development of audio-driven video generation tools with a focus on stability and efficiency.
RANK_REASON New open-source model release from a significant AI lab (Meituan).
Read on Hugging Face Trending Models →
- Hugging Face
- LongCat-Video-Avatar 1.5
- meituan-longcat
- Whisper-Large
- LongCat-Video
- Meituan
- meituan-longcat/LongCat-Video-Avatar-1.5
AI-generated summary · Google Gemini · from 5 sources. How we write summaries →