Researchers have developed Puppeteer, a novel diffusion model designed to generate co-speech gestures that are temporally coherent, semantically aligned with speech, and grounded with surrounding objects. Unlike previous models that primarily focused on audio-gesture alignment, Puppeteer explicitly incorporates posture constraints and object geometry. The model operates in a causal latent space, allowing for explicit temporal control and enabling tasks like gesture in-betweening and completion. To facilitate evaluation and development, a new synthetic dataset called SceneGes was created, featuring embodied co-speech gestures and corresponding 3D objects. AI
IMPACT This research advances generative AI capabilities in multimodal synthesis, potentially improving human-computer interaction and virtual character animation.
RANK_REASON The cluster describes a new research paper detailing a novel model for co-speech gesture generation. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →