Researchers have developed a new pretraining strategy for skeleton-based zero-shot spatio-temporal action localization, aiming to identify unseen actions in videos without extensive annotation. The method, called Skeleton-Language feature Pooling Switching, uses a weakly-supervised vision-language pretraining mechanism to align skeleton features with text embeddings. It also incorporates Scene-Mixed Discriminative Contrastive Learning within a MIL framework to differentiate actions at the instance level. Experiments on four datasets show the approach effectively mitigates annotation cost limitations. AI
IMPACT This research could improve the efficiency of video analysis systems by reducing the need for detailed action annotations.
RANK_REASON Academic paper detailing a new method for action localization. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- MIL framework
- Scene-Mixed Discriminative Contrastive Learning
- Skeleton-Language feature Pooling Switching
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →