VideoLLMs
PulseAugur coverage of VideoLLMs — every cluster mentioning VideoLLMs across labs, papers, and developer communities, ranked by signal.
1 day(s) with sentiment data
-
StateTrace framework enhances VideoLLM reasoning for invisible objects · 2 sources tracked
Researchers have developed StateTrace, a new object-centric framework designed to improve the spatiotemporal reasoning capabilities of VideoLLMs, particularly in scenarios involving long videos where objects may become …
-
New PoisonVID attack bypasses safety features in VideoLLMs
Researchers have developed a novel attack called PoisonVID that can bypass safety measures in Video Large Language Models (VideoLLMs). These models are used to moderate user-generated content by sampling key frames, but…
-
New GSTEP framework prunes VideoLLM tokens for efficiency
Researchers have developed GSTEP, a novel pruning framework designed to enhance the efficiency of Video Large Language Models (VideoLLMs). Unlike previous methods that prune tokens locally within segments, GSTEP models …
-
New dataset VideoNorms tests cultural awareness in VideoLLMs
Researchers have developed VideoNorms, a new dataset designed to evaluate the cultural awareness of Video Large Language Models (VideoLLMs). The dataset includes over 3,000 human judgments derived from popular US and Ch…
-
VideoLLMs can now watch and think simultaneously with new VST paradigm
Researchers have introduced Video Streaming Thinking (VST), a new paradigm designed to enable online Video Large Language Models (VideoLLMs) to process and reason about video content simultaneously. This approach aims t…
-
New benchmark NeMo challenges video LLMs on temporal understanding
Researchers have introduced NeMo, a novel task and benchmark called NeMoBench, designed to evaluate the temporal understanding capabilities of video large language models (VideoLLMs). The task, inspired by the 'needle i…