Video-MME
PulseAugur coverage of Video-MME — every cluster mentioning Video-MME across labs, papers, and developer communities, ranked by signal.
-
RT-NeuS framework accelerates video question answering with adaptive temporal verification
Researchers have developed RT-NeuS, a novel framework designed to significantly accelerate the process of long-form video question answering (LVQA). Traditional vision-language models (VLMs) struggle with the temporal c…
-
New benchmarks and architectures advance long-video understanding in MLLMs
Researchers are developing new methods to improve how multimodal large language models (MLLMs) understand long videos. One approach, MoTE, uses a Mixture of Task Experts to route computations to task-specific modules, e…
-
New frameworks tackle video reasoning challenges in large vision-language models
Researchers have developed new frameworks to address the challenges in video reasoning for large vision-language models (LVLMs). One approach, the Chain of Evidence (CoE), decouples grounding and reasoning to improve ef…
-
New MEDR method improves multimodal LLM video processing efficiency
Researchers have developed a new query-independent frame selection method called MEDR, designed to improve the efficiency of multimodal large language models when processing long videos. Unlike query-dependent methods t…
-
StreamOPD enhances streaming video understanding with new post-training technique
Researchers have introduced StreamOPD, a novel post-training technique designed to enhance streaming video understanding. This method utilizes on-policy distillation with verifiable rewards and a spatio-temporal cue-gat…
-
New VideoGAIA benchmark challenges advanced AI assistants in multi-turn video understanding
Researchers have introduced VideoGAIA, a new benchmark designed to evaluate the capabilities of multimodal large language models (MLLMs) in agentic video understanding. Unlike previous benchmarks that focus on single-tu…
-
New LAVE framework enhances video agent planning with latent visual evidence reuse
Researchers have introduced LAVE, a novel framework designed to enhance the planning capabilities of video tool-use agents. LAVE addresses the "Tool observation bottleneck" by enabling agents to reuse latent visual evid…
-
New frameworks boost AI long video understanding efficiency
Researchers have developed two new frameworks, EcoFrame and EviSelect, designed to improve the efficiency of long video understanding by large language models. EcoFrame uses a training-free approach that adapts the fram…
-
New GCR framework enhances long-video QA by optimizing frame selection
Researchers have introduced GCR, a novel framework designed to improve long-video question answering by optimizing the selection of relevant frames within a constrained budget. This training-free approach addresses limi…
-
FORGE method enhances LLM video understanding without retraining
Researchers have developed FORGE, a novel method for improving long-form video understanding in multimodal large language models (MLLMs). This model-agnostic technique operates at inference time without requiring additi…
-
New LENS framework enhances AI video understanding with adaptive keyframe sampling
Researchers have developed LENS, a novel framework designed to improve how Multi-modal Large Language Models (MLLMs) process long-form videos. LENS addresses the challenge of limited context windows by adaptively sampli…
-
New VideoTreeSearch framework enables self-correcting agents for long video QA
Researchers have introduced VideoTreeSearch (VTS), a novel framework designed to improve long-video question answering by treating the task as a self-correcting search over an adaptive temporal tree. Unlike previous met…
-
New method efficiently selects video frames for MLLM analysis
Researchers have developed a novel method called DAFS (Dynamic Attention-based Budget-aware Frame Selection) for efficiently selecting relevant frames from long videos for analysis by multimodal large language models (M…
-
New DELTAVID framework boosts video LLMs' fine-grained perception
Researchers have introduced DELTAVID, a novel framework designed to improve the fine-grained spatiotemporal perception capabilities of video multimodal large language models (Video MLLMs). This approach transforms the t…
-
New ReQuest pipeline enhances long-form video QA for LLMs
Researchers have developed ReQuest, a novel pipeline designed to improve question-answering capabilities for long-form videos. This method addresses the limitations of fixed input token budgets in multimodal large langu…
-
HiMu framework enhances long video question answering with hierarchical frame selection
Researchers have developed HiMu, a novel framework designed to improve frame selection for long-form video question answering tasks. This training-free system decomposes complex queries into a hierarchical logic tree, u…
-
InternVideo3 enhances video understanding with new reasoning framework
Researchers have introduced InternVideo3, a new framework designed to improve long-horizon video understanding and agentic capabilities. The system utilizes Multimodal Contextual Reasoning (MCR) to process video content…
-
ReTool-Video enhances video agents with recursive tool use
Researchers have introduced ReTool-Video, a novel approach for video understanding agents that enhances their reasoning capabilities. This method utilizes an expanded tool library with 134 specialized tools, including m…
-
LinMU achieves linear complexity for multimodal understanding models
Researchers have developed LinMU, a novel Vision-Language Model (VLM) architecture that achieves linear complexity, overcoming the quadratic complexity limitations of current models. This new design utilizes an M-MATE b…
-
Introducing gpt-realtime and Realtime API updates
OpenAI has released GPT-4.1, a new series of models for its API that offer significant improvements in coding, instruction following, and long context comprehension, outperforming previous models like GPT-4o. The compan…