Researchers have developed EMODY Flow, a novel framework for generating full-body motion that synchronizes with speech and emotional cues. This system addresses a limitation in existing models where emotion conditioning is often underutilized. EMODY Flow, a lightweight model with around 35 million parameters, integrates with a frozen Qwen-3 Omni model and utilizes its Mimi audio codecs. It employs two parallel Diffusion Transformer (DiT) generators for body pose and facial expressions, enhanced by an auxiliary emotion classifier to ensure distinct emotional expression in the generated motions. The framework achieves state-of-the-art results on the BEAT2 dataset for gesture quality, correlation, and diversity, significantly outperforming previous methods. AI
IMPACT Advances embodied AI by enabling more expressive and emotionally resonant character animations driven by audio.
RANK_REASON Academic paper detailing a new model and its performance on benchmarks. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →