Apple researchers have introduced LVSum, a new benchmark designed to evaluate how well multimodal large language models (MLLMs) can summarize long videos while maintaining temporal accuracy. The benchmark consists of 72 videos, averaging 16 minutes each, with human-generated summaries that include temporal references. Experiments using LVSum revealed that transcripts are more crucial than visual frames for summarization quality, and current MLLMs still struggle with temporal grounding and cross-modal coherence compared to human-written summaries. AI
IMPACT This benchmark could drive improvements in MLLMs for video understanding and summarization tasks.
RANK_REASON The cluster describes a new research paper introducing a benchmark for evaluating AI models.
Read on Apple Machine Learning Research →
- Alkesh B Patel
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Gotit.pub
- Hugging Face
- LVSum
- ScienceCast
- Apple Inc.
- European Conference on Computer Vision
- Ganesh Nagarajan
- Melis Ozyildirim
- multimodal large language models
- NarrativeTrack
- Ying-Chang Cheng
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →