PulseAugur
EN
LIVE 15:00:28

STARCaster video diffusion model enables identity-aware talking portraits

Researchers have developed STARCaster, a novel video diffusion model designed for creating talking portraits that are aware of both identity and viewpoint. This model integrates speech-driven animation and dynamic viewpoint control into a single framework, moving beyond the limitations of existing 2D models that rely heavily on reference guidance and 3D methods that can suffer from identity drift. STARCaster employs a compositional approach, starting with identity-aware motion, progressing to audio-visual synchronization, and finally enabling novel view animation. To address the scarcity of 4D audio-visual data, the model uses a decoupled learning strategy to train view consistency and temporal coherence independently. Evaluations indicate that STARCaster outperforms previous methods across various benchmarks for both tasks and identities. AI

IMPACT This model could advance realistic AI-driven avatar generation for applications like virtual assistants and content creation.

RANK_REASON The cluster contains a research paper detailing a new model. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CV →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

STARCaster video diffusion model enables identity-aware talking portraits

COVERAGE [1]

  1. arXiv cs.CV TIER_1 English(EN) · Foivos Paraperas Papantoniou, Stathis Galanakis, Rolandos Alexandros Potamias, Bernhard Kainz, Stefanos Zafeiriou ·

    STARCaster: Spatio-Temporal AutoRegressive Video Diffusion for Identity- and View-Aware Talking Portraits

    arXiv:2512.13247v2 Announce Type: replace Abstract: This paper presents STARCaster, an identity-aware spatio-temporal video diffusion model that addresses both speech-driven portrait animation and dynamic viewpoint control, given an identity embedding or reference image, within a…