Vision--Language Models
PulseAugur coverage of Vision--Language Models — every cluster mentioning Vision--Language Models across labs, papers, and developer communities, ranked by signal.
15 day(s) with sentiment data
-
VLMs struggle to reliably assess sidewalk accessibility, study finds
Researchers have investigated the capability of vision-language models (VLMs) to assess sidewalk accessibility attributes from pedestrian-level imagery. Using sampling-based conformal prediction, they evaluated four VLM…
-
New method deciphers how VLMs verbalize image semantics using OCR heads
Researchers have developed a method to understand how Vision-Language Models (VLMs) process image semantics, focusing specifically on their optical character recognition (OCR) capabilities. By identifying specific atten…
-
New benchmark reveals faithfulness gaps in Vision-Language Models
Researchers have introduced EDCT-Bench, a new benchmark designed to identify faithfulness issues in Vision-Language Models (VLMs). This benchmark uses an intervention-based protocol called Explanation-Driven Counterfact…
-
New framework enhances VLM safety with executable rules
Researchers have introduced GuardEn, a novel framework designed to enhance the safety of vision-language models (VLMs). This system decomposes complex safety policies into executable code, allowing for more adaptable an…
-
New benchmark evaluates vision-language models for disaster assessment
Researchers have introduced DisasterInsight, a new multimodal benchmark designed to evaluate vision-language models (VLMs) in disaster assessment. This benchmark focuses on building-centric analysis, going beyond genera…
-
New framework reveals hidden instability in Vision-Language Models
Researchers have identified a hidden instability in Vision-Language Models (VLMs) that is not captured by standard output-level assessments. A new evaluation framework measures internal embedding drift, spectral sensiti…
-
Robusto-2 paper benchmarks VLMs vs human drivers in autonomous driving
A new research paper, Robusto-2, benchmarks Vision-Language Models (VLMs) against human drivers in simulated autonomous driving scenarios. The study used dashcam footage from Lima and New York City, posing questions acr…
-
New method integrates trajectory planning into Vision-Language Models
Researchers have developed DiffAdapterVLA, a novel method that integrates continuous trajectory generation directly into the backbone of Vision--Language Models (VLMs). This approach injects explicit trajectory tokens i…
-
New methods accelerate Vision-Language Model inference by optimizing token processing · 2 sources tracked
Two new research papers propose methods to accelerate the inference of Vision-Language Models (VLMs) by reducing computational overhead. StackTok focuses on adaptive visual token selection, prioritizing query relevance …
-
New NoteVQA benchmark reveals VLM struggles with real-life visual questions
A new benchmark called NoteVQA has been developed to evaluate vision-language models (VLMs) on real-world visual questions, addressing the limitations of existing benchmarks that focus on predefined capabilities. The be…
-
VideoXAgent tackles long video understanding with online agent harness
Researchers have developed VideoXAgent, an online harness designed for understanding long videos. This system plans tasks, uses specialized tools like VLMs, OCR, and ASR on demand, and aggregates evidence to answer quer…
-
New system enables real-time video understanding with VLMs
Researchers have developed a novel system designed for real-time video understanding using Vision-Language Models (VLMs). This system integrates lightweight clients with a server runtime that handles speech recognition,…
-
New benchmark DementiaCare-Bench tests VLM capabilities in dementia care
Researchers have developed DementiaCare-Bench, a new video benchmark designed to evaluate the capabilities of video-language models (VLMs) in understanding and responding to behavioral and psychological symptoms of deme…
-
New FragileFlow Method Boosts Foundation Model Robustness
Researchers have introduced FragileFlow, a novel plug-in regularizer designed to enhance the robustness of foundation models, including LLMs and Vision-Language Models. This method addresses a failure mode where predict…
-
CS-CLIP enhances vision-language models for compositional reasoning
Researchers have developed CS-CLIP, a new approach to enhance vision-language models (VLMs) for compositional reasoning. Existing VLMs often show biases towards specific elements, leading to underperformance on complex …
-
Study: Canonical Color Decodable from Grayscale Images in VLMs
Researchers have explored how vision encoders within vision-language models (VLMs) represent conceptual information, specifically focusing on canonical colors. Their study demonstrates that even when color is removed fr…
-
New method improves OOD detection in medical AI using intermediate VLM layers
Researchers have developed a new method for out-of-distribution (OOD) detection in medical AI systems, addressing the challenge of domain shifts across different institutions and patient populations. Existing Vision-Lan…
-
New AeroBelief framework improves aerial object navigation for UAVs
Researchers have developed AeroBelief, a novel framework designed to enhance aerial object navigation for unmanned aerial vehicles (UAVs). This system addresses the challenges posed by noisy and transient visual data fr…
-
New CoVeR method prunes visual tokens for 3D reasoning in VLMs
Researchers have developed CoVeR, a novel method for pruning visual tokens in Vision-Language Models (VLMs) when processing 3D scenes represented by multi-view images. This technique addresses the issue of redundant tok…
-
New benchmark reveals vision-language models struggle with personalized safety
Researchers have introduced MPS-Bench, a new benchmark designed to evaluate personalized safety in vision-language models (VLMs). The benchmark consists of 5,181 scenarios derived from real-world images and includes hid…