Researchers have introduced "Identity-as-Presence," a novel framework designed for personalized audio-video generation that can handle multiple subjects simultaneously. This system addresses limitations in current methods, which often focus on single individuals and struggle with precise visual and vocal identity alignment in multi-subject scenarios. The framework utilizes an automated data curation pipeline to create labeled audio-visual pairs and a unified injection mechanism for binding appearance and voice through shared cross-modal identity binding and subject-anchored captions. Experiments indicate that Identity-as-Presence achieves superior audio quality, video fidelity, and audio-visual consistency, with improved multi-subject binding compared to existing approaches. AI
IMPACT This research advances personalized content creation by enabling more sophisticated multi-subject audio-visual synthesis.
RANK_REASON The cluster contains a new academic paper detailing a novel framework for audio-video generation. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- Identity-as-Presence
- ScienceCast
- Yingjie Chen
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →