MVBench
PulseAugur coverage of MVBench — every cluster mentioning MVBench across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New GAM-Agent framework boosts visual reasoning in LLMs via game theory
Researchers have developed GAM-Agent, a novel framework that enhances visual reasoning in large language models by employing a game-theoretic approach. This system treats the reasoning process as a non-zero-sum game whe…
-
New benchmark diagnoses visual grounding in video LLMs
A new paper introduces the Visual Dependency Gap (VDG) to assess the visual grounding capabilities of video large language models (LLMs). The VDG measures the difference in accuracy between models processing original vi…
-
Goal-Driven Data Optimization speeds up multimodal AI training
Researchers have developed a framework called Goal-Driven Data Optimization (GDO) to improve the efficiency of multimodal instruction tuning. GDO computes sample descriptors to create optimized training subsets tailored…
-
VisReflect framework improves LVLM fine-grained perception in long contexts
Researchers have introduced VisReflect, a novel framework designed to enhance fine-grained perception in Large Vision Language Models (LVLMs) when processing high-resolution images and long videos. This method addresses…
-
ReTool-Video enhances video agents with recursive tool use
Researchers have introduced ReTool-Video, a novel approach for video understanding agents that enhances their reasoning capabilities. This method utilizes an expanded tool library with 134 specialized tools, including m…
-
VideoThinker framework improves lightweight MLLMs' video reasoning via causal debiasing
Researchers have developed VideoThinker, a novel framework designed to enhance the reasoning capabilities of lightweight multimodal language models (MLLMs) in video analysis. This approach addresses the issue of percept…
-
ReGATE method accelerates multimodal LLM training by selectively pruning tokens
Researchers have developed ReGATE, a novel method to accelerate the training of multimodal large language models (MLLMs) by adaptively pruning tokens. This technique uses a teacher-student framework where a frozen teach…
-
New PushupBench benchmark reveals VLMs struggle with counting repetitions
Researchers have introduced PushupBench, a new dataset designed to evaluate the ability of vision-language models (VLMs) to accurately count repetitions in videos. The benchmark highlights that even top-tier VLMs strugg…