LongVideoBench
PulseAugur coverage of LongVideoBench — every cluster mentioning LongVideoBench across labs, papers, and developer communities, ranked by signal.
5 day(s) with sentiment data
-
New LAVE framework enhances video agent planning with latent visual evidence reuse
Researchers have introduced LAVE, a novel framework designed to enhance the planning capabilities of video tool-use agents. LAVE addresses the "Tool observation bottleneck" by enabling agents to reuse latent visual evid…
-
New frameworks boost AI long video understanding efficiency
Researchers have developed two new frameworks, EcoFrame and EviSelect, designed to improve the efficiency of long video understanding by large language models. EcoFrame uses a training-free approach that adapts the fram…
-
New GCR framework enhances long-video QA by optimizing frame selection
Researchers have introduced GCR, a novel framework designed to improve long-video question answering by optimizing the selection of relevant frames within a constrained budget. This training-free approach addresses limi…
-
FORGE method enhances LLM video understanding without retraining
Researchers have developed FORGE, a novel method for improving long-form video understanding in multimodal large language models (MLLMs). This model-agnostic technique operates at inference time without requiring additi…
-
New frameworks boost MLLM long-video understanding by adaptive frame processing · 3 sources tracked
Three new research papers introduce novel frameworks for enhancing the long-video understanding capabilities of multimodal large language models (MLLMs). These approaches aim to overcome the limitations of fixed context…
-
New VIBE method improves video-to-text model summaries without annotations
Researchers have developed VIBE, an annotation-free evaluation method for video-to-text models. VIBE assesses summaries based on their grounding in visual content and their utility for downstream tasks, aiming to overco…
-
New DELTAVID framework boosts video LLMs' fine-grained perception
Researchers have introduced DELTAVID, a novel framework designed to improve the fine-grained spatiotemporal perception capabilities of video multimodal large language models (Video MLLMs). This approach transforms the t…
-
New ReQuest pipeline enhances long-form video QA for LLMs
Researchers have developed ReQuest, a novel pipeline designed to improve question-answering capabilities for long-form videos. This method addresses the limitations of fixed input token budgets in multimodal large langu…
-
New QCA framework enhances long video understanding by optimizing keyframe selection
Researchers have developed a new framework called QCA for selecting keyframes in long videos to improve video understanding. This method is query- and content-aware, meaning it prioritizes frames that are relevant to a …
-
New STAR Framework Boosts LLM Video Analysis Capabilities
Researchers have developed a Spatiotemporal Reasoning Framework (STAR) to enhance the video question answering capabilities of multimodal large language models (MLLMs). STAR equips models like GPT-4o with a Video Toolki…
-
HiMu framework enhances long video question answering with hierarchical frame selection
Researchers have developed HiMu, a novel framework designed to improve frame selection for long-form video question answering tasks. This training-free system decomposes complex queries into a hierarchical logic tree, u…
-
Reflect-R1 framework improves AI video understanding with evidence-driven self-correction
Researchers have introduced Reflect-R1, a novel framework designed to enhance self-correction in long video understanding models. This system addresses the issue of models becoming overconfident due to a lack of externa…
-
Kwai releases Keye-VL-2.0 for long-video understanding
Kwai has released Keye-VL-2.0-30B-A3B, an open-source multimodal foundation model designed for long-video understanding and agentic intelligence. This model utilizes DeepSeek Sparse Attention to process up to 256K conte…
-
CREST method efficiently selects key frames from long videos
Researchers have developed CREST, a novel method for efficiently selecting key frames from long videos. This training-free approach leverages the temporal geometry of query-frame relevance, specifically focusing on loca…
-
GridProbe cuts VLM compute cost for long videos
Researchers have developed GridProbe, a novel method to improve the efficiency of long-video Visual Language Models (VLMs). This technique adaptively selects relevant frames during inference, reducing the computational …
-
LinMU achieves linear complexity for multimodal understanding models
Researchers have developed LinMU, a novel Vision-Language Model (VLM) architecture that achieves linear complexity, overcoming the quadratic complexity limitations of current models. This new design utilizes an M-MATE b…