Researchers have introduced MolmoMotion, a novel model designed for language-guided 3D motion forecasting. This model predicts the future 3D trajectories of points on an object based on a video frame, initial point positions, and a textual description of the intended action. MolmoMotion significantly outperforms existing methods on a new benchmark, PointMotionBench, and demonstrates utility in downstream applications like robotics and video generation. The project also includes MolmoMotion-1M, a large dataset of 3D trajectories paired with action descriptions, derived from over a million videos. AI
IMPACT Enhances capabilities in robotics and generative AI by enabling more accurate prediction of object motion based on natural language commands.
RANK_REASON Research paper and model release detailing a new approach to 3D motion forecasting.
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →