Researchers have developed a new framework called \"methodname\" for training-free composed video retrieval (CoVR). This approach leverages frozen foundation models by adapting their inference depth based on query difficulty. The framework first uses compact, reusable video-only representations for initial searches, then employs bounded reranking and candidate expansion for uncertain queries, and finally uses multimodal verification for close candidates. This adaptive strategy allows for scalable retrieval with fine-grained reasoning without requiring task-specific training, achieving state-of-the-art performance on benchmarks like Dense-WebVid-CoVR and CoVR-R. AI
IMPACT This framework could enable more efficient and scalable video search by adaptively using foundation models.
RANK_REASON This is a research paper detailing a new framework for video retrieval. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Composed Video Retrieval
- Connected Papers
- DagsHub
- Dense-WebVid-CoVR
- foundation model
- Gotit.pub
- Hugging Face
- Litmaps
- ScienceCast
- scite Smart Citations
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →