LVBench
PulseAugur coverage of LVBench — every cluster mentioning LVBench across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
VideoXAgent tackles long video understanding with online agent harness
Researchers have developed VideoXAgent, an online harness designed for understanding long videos. This system plans tasks, uses specialized tools like VLMs, OCR, and ASR on demand, and aggregates evidence to answer quer…
-
New methods enhance AI's understanding of long videos using synthetic data and temporal analysis · 4 sources tracked
Researchers are developing new methods to improve how large multimodal models understand long videos. One approach, SynMulti, uses a synthetic data generation pipeline to create unlimited annotated video data for tasks …
-
New benchmarks and architectures advance long-video understanding in MLLMs
Researchers are developing new methods to improve how multimodal large language models (MLLMs) understand long videos. One approach, MoTE, uses a Mixture of Task Experts to route computations to task-specific modules, e…
-
New CARVE probe audits video agent reliability by destroying frame content
Researchers have developed CARVE, a new black-box counterfactual probe designed to audit the reliability of video agents. CARVE compares how an agent's answer changes when the semantic content of retrieved video frames …
-
MERIT framework simplifies ultra-long video understanding
Researchers have developed MERIT, a novel framework for understanding ultra-long videos that exceed practical processing limits for current multi-modal large language models. MERIT employs a two-stage approach, first co…
-
New frameworks boost MLLM long-video understanding by adaptive frame processing · 3 sources tracked
Three new research papers introduce novel frameworks for enhancing the long-video understanding capabilities of multimodal large language models (MLLMs). These approaches aim to overcome the limitations of fixed context…
-
New VideoTreeSearch framework enables self-correcting agents for long video QA
Researchers have introduced VideoTreeSearch (VTS), a novel framework designed to improve long-video question answering by treating the task as a self-correcting search over an adaptive temporal tree. Unlike previous met…
-
Goal-Driven Data Optimization speeds up multimodal AI training
Researchers have developed a framework called Goal-Driven Data Optimization (GDO) to improve the efficiency of multimodal instruction tuning. GDO computes sample descriptors to create optimized training subsets tailored…
-
New DELTAVID framework boosts video LLMs' fine-grained perception
Researchers have introduced DELTAVID, a novel framework designed to improve the fine-grained spatiotemporal perception capabilities of video multimodal large language models (Video MLLMs). This approach transforms the t…
-
OmniAgent uses active perception for efficient video understanding · 2 sources tracked
Researchers have introduced OmniAgent, a novel omni-modal agent designed for video understanding that utilizes an iterative Observation-Thought-Action cycle based on Partially Observable Markov Decision Processes (POMDP…
-
New benchmarks and models advance vision-language capabilities in robotics and reasoning · 10 sources tracked
Recent research explores advancements in vision-language models (VLMs) across several domains. DeCAL introduces a new model for dexterous manipulation that integrates tactile sensing and visual-language understanding. R…