Vision--Language Models
PulseAugur coverage of Vision--Language Models — every cluster mentioning Vision--Language Models across labs, papers, and developer communities, ranked by signal.
24 day(s) with sentiment data
-
New VLM backdoor allows arbitrary, programmable control
Researchers have developed a novel method for implanting programmable backdoors into Vision-Language Models (VLMs). Unlike previous static backdoor attacks, this new technique allows attackers to dynamically control tar…
-
New methods for efficient visual token compression in VLMs unveiled
Two new research papers propose methods for compressing visual tokens in vision-language models (VLMs) to improve efficiency. The first, "Not All Visual Tokens Are Equally Safe to Remove," introduces a consequence-sensi…
-
New benchmark and dataset evaluate VLM capabilities in diagnosing tomato leaf diseases
Researchers have introduced TomaMMU, a large-scale dataset for understanding tomato leaf diseases, and TomaBench, a benchmark designed to evaluate Vision-Language Models (VLMs) on this task. The dataset includes over 28…
-
Onboard VLMs enable bandwidth-efficient Earth observation via dialogue
Researchers have developed a novel "Summarize First, Download Later" paradigm for Earth observation satellites, utilizing onboard Vision-Language Models (VLMs) to address bandwidth limitations. This approach involves th…
-
New benchmark DATAREEL reveals VLM struggles with automated video story generation
Researchers have introduced DATAREEL, a new benchmark designed to evaluate the capabilities of vision-language models (VLMs) in automatically generating data-driven video stories. The benchmark consists of 328 real-worl…
-
New SABRE framework automates stress testing for vision-language models
Researchers have developed SABRE, a new automated pipeline designed to create stress tests for vision-language models (VLMs). This framework converts task designs into structured specifications, images, and question-ans…
-
New WNM-3D model enhances 3D scene conditioning for navigation
Researchers have introduced WNM-3D, a novel World Navigation Model that incorporates 3D scene conditioning for closed-loop vision-language navigation (VLN). This model addresses limitations in current VLN systems by exp…
-
New benchmark VQABench analyzes cost-quality of cloud VLM VQA systems
A new benchmark called VQABench has been developed to evaluate the cost-quality trade-offs of cloud-based Vision-Language Models (VLMs) for visual question answering (VQA) systems. The research highlights that client-si…
-
New frameworks and benchmarks advance video anomaly detection capabilities
Researchers are developing advanced methods for video anomaly detection (VAD), a critical task for industrial applications and safety systems. New frameworks like VTO and FedVAR aim to improve generalization and address…
-
BIM-Native Tokenization Enhances Room Layout Synthesis
Researchers have developed a novel BIM-native tokenization method for synthesizing room layouts within Building Information Modeling (BIM) scenes. This approach encodes each room as a sequence of BIM-Token Bundles, unif…
-
New research explores visual evidence representation for AI item difficulty prediction
Researchers have explored methods for representing visual evidence in item difficulty prediction, comparing visual textualization (expressing images in language) with image-native modeling (retaining the original image)…
-
New GST-Bench benchmark reveals VLM struggles with global spatial awareness in video
Researchers have introduced GST-Bench, a new benchmark designed to evaluate the global spatial awareness of Vision-Language Models (VLMs) using video data. The benchmark, which includes questions derived from over 6,790…
-
New CROSS method enhances remote sensing image segmentation
Researchers have developed a new method called CROSS for referring remote sensing image segmentation. This approach aims to address limitations in existing Vision-Language Models (VLMs) and the Segment Anything Model (S…
-
New method enhances 3D indoor scene generation using graph validation
Researchers have developed a new method called Global Graph-Validated Optimization for generating 3D indoor scenes from text instructions. This approach uses a graph-based representation to separate semantic coherence f…
-
New BanglaWild benchmark evaluates Bengali scene text recognition for OCR and VLMs
Researchers have introduced BanglaWild, a new benchmark designed to evaluate Bengali scene text recognition for both optical character recognition (OCR) systems and vision-language models (VLMs). The benchmark consists …
-
New benchmarks and methods advance multimodal reasoning in AI
Researchers are developing new methods for multimodal knowledge graph completion and reasoning, integrating vision-language models (VLMs) with graph structures. ViSR-KGC proposes a visual subgraph reasoning approach tha…
-
UniEvo-RS framework enhances remote sensing segmentation with exemplar-driven prototypes
Researchers have developed UniEvo-RS, a novel framework for remote sensing segmentation that utilizes an omni-prompt approach with representative exemplar-driven prototype evolution. This system aims to overcome the per…
-
LLMs and VLMs show sensitivity to text casing, influencing attention
A new research paper explores how Large Language Models (LLMs) and Vision-Language Models (VLMs) are sensitive to letter casing, similar to human visual perception. The study found that formatting text in uppercase or a…
-
New XSPA attack method targets Vision-Language Models
Researchers have developed XSPA, a novel method for creating adversarial perturbations on Vision-Language Models (VLMs). This technique crafts imperceptible X-shaped sparse perturbations that can significantly degrade V…
-
New VC-Tooler framework enhances visual tool use for Vision--Language Models
Researchers have introduced VC-Tooler, a new framework designed to enhance the capabilities of Vision--Language Models (VLMs) in utilizing visual tools. Unlike previous methods that focused on single-tool grounding, VC-…