Researchers have introduced MELON, a new large-scale dataset designed for multi-event text-to-long-video retrieval, addressing the limitations of existing datasets that primarily focus on short clips with single events. MELON annotates multiple event intervals within long videos, enabling models to better understand complex, multi-event structures. Concurrently, a framework named TAME has been proposed, which enhances text-video retrieval by integrating temporal modeling into CLIP-based architectures. TAME utilizes Mixture-of-Experts layers and Frame-Temporal tokens to capture both local frame details and long-range temporal dependencies, showing improved performance on standard retrieval benchmarks. AI
IMPACT Advances in text-video retrieval could lead to more sophisticated content search and recommendation systems.
RANK_REASON Two research papers introducing new datasets and frameworks for text-video retrieval.
Read on arXiv cs.IR (Information Retrieval) →
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →