Researchers have developed ReQuest, a novel pipeline designed to improve question-answering capabilities for long-form videos. This method addresses the limitations of fixed input token budgets in multimodal large language models by employing an uncertainty-driven, question-adaptive keyframe selection process. ReQuest integrates a lightweight selector, a routing mechanism that triggers additional inference based on model uncertainty, and an adaptive non-maximum suppression technique to select relevant and temporally diverse frames. The system functions as a plug-and-play solution, enhancing performance on benchmarks like Video-MME, MLVU, and LongVideoBench without altering the core MLLM. AI
IMPACT This method could improve the efficiency and accuracy of AI systems processing long video content for question-answering tasks.
RANK_REASON The cluster describes a new research paper detailing a novel method for improving video question-answering.
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →