Video-Language Models
PulseAugur coverage of Video-Language Models — every cluster mentioning Video-Language Models across labs, papers, and developer communities, ranked by signal.
3 day(s) with sentiment data
-
New benchmark DrivelHub+ tests AI's grasp of social media video nuance
A new benchmark called DrivelHub+ has been introduced to evaluate the ability of video-language models to understand implicit and non-literal meanings in social media videos. This benchmark consists of 1,000 annotated v…
-
New benchmark evaluates video-language models on long-form descriptions
Researchers have introduced CLIP-CC-Bench, a new evaluation suite designed to assess the capabilities of video-language models in generating detailed, paragraph-length descriptions of video content. This benchmark, deri…
-
TOPReward uses VLM token probabilities for robot learning rewards
Researchers have developed TOPReward, a novel method for generating dense, instruction-conditioned feedback for robotic learning without requiring manual annotations or task-specific reward models. This approach leverag…
-
PhysMRV framework enhances VLM physical reasoning without training
Researchers have developed PhysMRV, a novel framework designed to enhance the physical plausibility reasoning capabilities of video-language models (VLMs). This training-free approach transforms videos into a structured…
-
New benchmark and method improve text-to-video retrieval for ecological data
Researchers have introduced Prompting-MammAlps, a new benchmark for fine-grained text-to-video retrieval specifically designed for camera-trap data. This benchmark aims to address the limitations of current video-langua…
-
New CycliST Benchmark Tests Video Language Models on Cyclical Reasoning
Researchers have introduced CycliST, a new benchmark dataset designed to test the capabilities of Video Language Models (VLMs) in understanding and reasoning about cyclical state transitions. The dataset features synthe…
-
New framework enhances Video-Language Models with flexible text input
Researchers have proposed a new framework to improve Video-Language Models (VLMs) by addressing limitations in text input. Current VLMs often rely on predefined text templates, which are restrictive and time-consuming t…