Video Large Language Models
PulseAugur coverage of Video Large Language Models — every cluster mentioning Video Large Language Models across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
New SLVMBench benchmark reveals video LLMs struggle with skill learning from long memory
Researchers have introduced SLVMBench, a novel benchmark designed to evaluate the ability of video large language models (video-LLMs) to learn skills from extended video memory and apply them in real-time scenarios. The…
-
GeoTrace framework compresses video tokens for efficient Video LLMs
Researchers have introduced GeoTrace, a novel framework designed to enhance the efficiency of Video Large Language Models (Video LLMs) by compressing visual tokens. This training-free method decomposes video evidence in…
-
New benchmarks push video AI to ground answers in temporal evidence · 4 sources tracked
Two new research papers introduce benchmarks and models for video question answering that focus on temporal reasoning and evidence grounding. The EG-VQA benchmark, with over 11,000 QA pairs and temporal evidence annotat…
-
VideoLLMs exhibit 'bag-of-events' behavior, hallucinating temporal links
A new study published on arXiv introduces DistractionBench, a framework designed to test the temporal understanding capabilities of Video Large Language Models (VideoLLMs). Researchers found that these models often exhi…
-
EvoVid framework enables Video-LLMs to self-evolve using raw video data
Researchers have introduced EvoVid, a novel framework designed to enhance Video Large Language Models (Video-LLMs) through temporal-centric self-evolution. Unlike previous self-evolving methods that are limited to stati…