Two new research papers address the challenge of enabling multimodal large language models (MLLMs) to understand long videos, which is currently limited by token and computational budgets. The first paper, "Evaluation of MLLM-Agnostic Plug-and-Play Keyframe Selection Methods for Long Video Understanding," provides a comprehensive evaluation of existing training-free keyframe selection methods, finding that QAaF generally performs best. The second paper introduces MarKey, a novel training-free framework that uses a greedy optimization approach to select keyframes by considering query relevance, marginal coverage gain, and context-dependent redundancy, demonstrating superior performance across multiple benchmarks and MLLM backbones. AI
IMPACT These methods could significantly improve the efficiency and accuracy of AI systems processing long video content, enabling new applications in analysis and summarization.
RANK_REASON Two academic papers published on arXiv introducing new methods for video understanding with LLMs.
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- FOCUS
- Gotit.pub
- Hugging Face
- MarKey
- MLLMs
- multimodal large language model
- ScienceCast
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →