Researchers have introduced GCR, a novel framework designed to improve long-video question answering by optimizing the selection of relevant frames within a constrained budget. This training-free approach addresses limitations in existing methods by preventing frame clustering and enhancing the alignment between textual and visual evidence. GCR achieves this through a three-stage process: Ground, Cover, and Refine, which jointly curates evidence by selecting query-relevant anchors, supplementing them with complementary visual information, and revisiting omitted regions for potentially greater value. Experiments on benchmarks like LongVideoBench and Video-MME show consistent improvements across various backbones and frame budgets. AI
IMPACT This framework could improve the efficiency and accuracy of AI systems processing long video content.
RANK_REASON The cluster contains a research paper detailing a new framework for video question answering. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →