Researchers have developed STARCaster, a novel video diffusion model designed for creating talking portraits that are aware of both identity and viewpoint. This model integrates speech-driven animation and dynamic viewpoint control into a single framework, moving beyond the limitations of existing 2D models that rely heavily on reference guidance and 3D methods that can suffer from identity drift. STARCaster employs a compositional approach, starting with identity-aware motion, progressing to audio-visual synchronization, and finally enabling novel view animation. To address the scarcity of 4D audio-visual data, the model uses a decoupled learning strategy to train view consistency and temporal coherence independently. Evaluations indicate that STARCaster outperforms previous methods across various benchmarks for both tasks and identities. AI
IMPACT This model could advance realistic AI-driven avatar generation for applications like virtual assistants and content creation.
RANK_REASON The cluster contains a research paper detailing a new model. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Foivos Paraperas Papantoniou
- Gotit.pub
- Hugging Face
- ScienceCast
- STARCaster
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →