Researchers have introduced AnyMo, a novel framework for generating human motion conditioned on various modalities like text, speech, and music. This approach utilizes a masked modeling transformer and a motion tokenizer, trained on the newly created OmniHuMo dataset, which contains over 5,000 hours of motion data with multimodal annotations. AnyMo aims to overcome limitations of previous methods by enabling flexible control and high-fidelity synthesis across arbitrary combinations of input signals. AI
IMPACT Enables more flexible and high-fidelity human motion generation from diverse multimodal inputs.
RANK_REASON This is a research paper describing a new model and dataset for motion generation.
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →