Researchers have developed MoVT, a new framework designed to improve text-to-3D human motion generation by utilizing extensive human action video data. The core of MoVT is a cross-modal augmented motion tokenizer that projects 3D motion tokens into the 2D domain, enriching the motion codebook with real-world patterns from videos. These enhanced codebooks are then integrated into a generative masked transformer, enabling modality-agnostic prediction of motion token indices. This approach allows for the use of text-index pairs derived from 2D codebooks and annotated motion videos to further refine the generator, outperforming existing state-of-the-art methods in empirical evaluations. AI
IMPACT This research could lead to more sophisticated and naturalistic 3D character animations driven by text prompts.
RANK_REASON The cluster contains a research paper detailing a new framework for text-to-motion generation. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →