This paper explores three distinct methods for generating videos with multiple subjects from a single image and text prompt: direct, parallel, and sequential generation. Direct generation attempts to synthesize all subjects and motions simultaneously, while parallel generation creates each subject independently before composing them, potentially sacrificing inter-subject context. Sequential generation builds the video progressively, introducing subjects one by one, which can preserve scene context but is sensitive to ordering and error propagation. The study evaluates these approaches based on appearance preservation, motion fidelity, temporal consistency, and inter-subject coherence, offering insights into their respective strengths and weaknesses for controllable multi-subject video generation. AI
IMPACT Provides insights into designing controllable multi-subject video generation systems.
RANK_REASON The cluster contains an academic paper detailing a comparative study of different methods for a specific AI task.
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →