Researchers have developed MOCO, a novel diffusion-based framework for generating coherent 3D avatar motions from multiple simultaneous inputs like speech audio, text descriptions, and trajectory data. Unlike previous methods that struggle with aligned multimodal data and often produce mismatched movements, MOCO decouples the motion generation process. In each denoising step, it independently generates motions for each modality and then assembles them based on spatial rules, iteratively refining the overall motion for natural and fluid results. Experiments on a custom benchmark show MOCO outperforms existing approaches in multimodal motion generation. AI
IMPACT This research could lead to more realistic and interactive 3D avatars in gaming, virtual reality, and animation.
RANK_REASON The cluster contains a research paper detailing a new method for multimodal motion generation. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Connected Papers
- CORE Recommender
- DagsHub
- Gotit.pub
- Hugging Face
- Litmaps
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →