New benchmarks and architectures advance long-video understanding in MLLMs
ByPulseAugur Editorial·[13 sources]·
Researchers are developing new methods to improve how multimodal large language models (MLLMs) understand long videos. One approach, MoTE, uses a Mixture of Task Experts to route computations to task-specific modules, enhancing accuracy and efficiency in video-language tasks. Another area of focus is optimizing how models extract relevant information from lengthy videos, with techniques like VideoHarness-RSI recursively searching for effective context-construction programs and TRACE focusing on evidence-supported answers. Additionally, Video-IFBench has been introduced as a benchmark to evaluate MLLMs' ability to follow diverse instructions within video understanding scenarios, highlighting current challenges in constraint adherence.
AI
IMPACT
Advances in long-video understanding could enable more sophisticated AI agents and analysis tools for complex visual data.
RANK_REASON
Multiple research papers introducing new architectures, benchmarks, and methods for long-video understanding.
Long-video MLLMs must model temporal change before a limited visual-token budget removes most frame evidence. We introduce LongVU-TTT, which inserts a convolutional Test-Time Training (TTT) resampler with causal fast-weight updates between the vision encoder and the LLM. Its grou…
Vision-language models (VLMs) have made substantial progress in long-video understanding, with standard backbone models typically answering questions from frames sampled across the full video. However, as videos become longer, the full-video context inevitably contains more quest…
arXiv:2608.24302v1 Announce Type: new Abstract: Long-video understanding depends critically on how a limited model context is constructed from a much longer video. Existing approaches improve this process through compression, retrieval, memory, and agentic evidence acquisition, b…
arXiv cs.LG
TIER_1English(EN)·Muhammad Asad Ali, Umar Khan, Nadia Robertini, Didier Stricker·
arXiv:2608.24763v1 Announce Type: cross Abstract: Procedural video-language models must solve heterogeneous tasks from the same visual evidence, including action recognition, forecasting, and procedure prediction. Dense transformer decoders share the same feed-forward networks ac…
arXiv:2608.22516v1 Announce Type: cross Abstract: A long-video answer is evidence-supported only when the frames decoded from the video cover every event the answer depends on. Existing evaluations score final-answer correctness or predicted evidence intervals, but the frames a m…
arXiv:2608.27065v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) has recently emerged as an effective post-training paradigm that improves policy optimization through dense token-level supervision from a privileged self-teacher. Despite its promise, OPSD remains…
arXiv cs.CV
TIER_1English(EN)·Kaishen Wang, Dongdi Zhao, Yijun Liang, Dingqiang Ye, Ruibo Chen, Heng Huang, Di Fu·
arXiv:2608.25356v1 Announce Type: new Abstract: Vision-language models (VLMs) have made substantial progress in long-video understanding, with standard backbone models typically answering questions from frames sampled across the full video. However, as videos become longer, the f…
arXiv:2608.25529v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have shown strong performance in video understanding. However, their ability to follow instructions in this domain remains under-explored. Real-world video understanding requires models not o…
arXiv:2608.25729v1 Announce Type: new Abstract: Long-video MLLMs must model temporal change before a limited visual-token budget removes most frame evidence. We introduce LongVU-TTT, which inserts a convolutional Test-Time Training (TTT) resampler with causal fast-weight updates …
arXiv:2608.20805v1 Announce Type: new Abstract: Long-form video understanding remains challenging for video agents due to the mismatch between query demands and evidence acquisition strategies. Although recent planning-before-perception methods outperform query-agnostic pipelines…
arXiv cs.CV
TIER_1English(EN)·Beibei Zhang, Chao Xu, Jun Lan, Zongyi Li, Lai Wei, Huijia Zhu, Tongwei Ren·
arXiv:2608.20814v1 Announce Type: new Abstract: Though Multimodal Large Language Models (MLLMs) have shown impressive potential in video understanding, long video understanding (LVU) remains challenging since distracting noise in complex and lengthy contexts can obscure localized…