VideoQA
PulseAugur coverage of VideoQA — every cluster mentioning VideoQA across labs, papers, and developer communities, ranked by signal.
2 day(s) with sentiment data
-
New CapMem benchmark tests caption-based memory for egocentric video
Researchers have introduced CapMem, a new benchmark designed to evaluate episodic memory capabilities in egocentric videos for wearable assistants. The benchmark, comprising 75 videos and 1,000 questions, explores wheth…
-
SVMemAgent tackles online frame selection for streaming video
Researchers have developed SVMemAgent, a novel system designed for online frame selection in streaming video scenarios. Unlike traditional methods that require full video and query access beforehand, SVMemAgent operates…
-
New Parallel Tube Decoding method slashes video grounding latency
Researchers have developed a new method called Parallel Tube Decoding (PTD) to improve the efficiency and accuracy of spatio-temporal video grounding. This technique removes autoregressive dependencies, significantly re…
-
New research optimizes visual token processing for long-video MLLMs
Researchers are exploring methods to optimize how multimodal large language models (MLLMs) process visual information, particularly for long videos. Several papers introduce techniques for selecting, compressing, and pr…
-
New datasets and methods advance causal reasoning in video question answering · 2 sources tracked
Researchers have introduced two new datasets and methodologies for causal video question answering, aiming to improve models' ability to understand complex cause-and-effect relationships in dynamic visual scenes. Causal…
-
New algorithm TASKER improves video understanding and agentic tasks
Researchers have developed TASKER, a novel keyframe extraction algorithm designed to improve performance in both Video Question Answering (VideoQA) and video-guided agentic tasks. This algorithm, detailed in a new paper…
-
New framework uses counterfactual reasoning to improve video QA systems
Researchers have developed a new framework called CREDiT to improve the reliability of video question-answering systems. This framework uses counterfactual reasoning and structural causal models to disentangle causal ev…
-
New EBM-RL framework enhances video role-playing with visual grounding
Researchers have developed a new framework called EBM-RL, which uses a decoupled approach to improve role-playing dialogue in immersive video applications. This method explicitly separates visual perception, reasoning, …
-
CurEvo framework enhances video understanding via curriculum-guided self-evolution
Researchers have introduced CurEvo, a novel framework designed to enhance self-evolutionary video understanding models. This approach integrates curriculum learning to provide structured guidance, addressing limitations…
-
EMCompress introduces novel compression for Video-LLMs, improving efficiency
Researchers have introduced EMCompress, a novel method for improving the efficiency of Video-LLMs in long-video reasoning tasks. This approach uses a cognitively-inspired technique called Endomorphic Multimodal Compress…