PulseAugur
EN
LIVE 19:16:22
ENTITY Vision--Language Models

Vision--Language Models

PulseAugur coverage of Vision--Language Models — every cluster mentioning Vision--Language Models across labs, papers, and developer communities, ranked by signal.

Show in brief
Total · 30d
68
181 over 90d
Releases · 30d
0
0 over 90d
Papers · 30d
66
174 over 90d
TIER MIX · 90D
TOPICS
RELATIONSHIPS
SENTIMENT · 30D

24 day(s) with sentiment data

RECENT · PAGE 1/10 · 181 TOTAL
  1. TOOL · CL_196221 ·

    New VLM backdoor allows arbitrary, programmable control

    Researchers have developed a novel method for implanting programmable backdoors into Vision-Language Models (VLMs). Unlike previous static backdoor attacks, this new technique allows attackers to dynamically control tar…

  2. RESEARCH · CL_193544 ·

    New methods for efficient visual token compression in VLMs unveiled

    Two new research papers propose methods for compressing visual tokens in vision-language models (VLMs) to improve efficiency. The first, "Not All Visual Tokens Are Equally Safe to Remove," introduces a consequence-sensi…

  3. TOOL · CL_193515 ·

    New benchmark and dataset evaluate VLM capabilities in diagnosing tomato leaf diseases

    Researchers have introduced TomaMMU, a large-scale dataset for understanding tomato leaf diseases, and TomaBench, a benchmark designed to evaluate Vision-Language Models (VLMs) on this task. The dataset includes over 28…

  4. TOOL · CL_191424 ·

    Onboard VLMs enable bandwidth-efficient Earth observation via dialogue

    Researchers have developed a novel "Summarize First, Download Later" paradigm for Earth observation satellites, utilizing onboard Vision-Language Models (VLMs) to address bandwidth limitations. This approach involves th…

  5. TOOL · CL_191250 ·

    New benchmark DATAREEL reveals VLM struggles with automated video story generation

    Researchers have introduced DATAREEL, a new benchmark designed to evaluate the capabilities of vision-language models (VLMs) in automatically generating data-driven video stories. The benchmark consists of 328 real-worl…

  6. TOOL · CL_191242 ·

    New SABRE framework automates stress testing for vision-language models

    Researchers have developed SABRE, a new automated pipeline designed to create stress tests for vision-language models (VLMs). This framework converts task designs into structured specifications, images, and question-ans…

  7. TOOL · CL_191149 ·

    New WNM-3D model enhances 3D scene conditioning for navigation

    Researchers have introduced WNM-3D, a novel World Navigation Model that incorporates 3D scene conditioning for closed-loop vision-language navigation (VLN). This model addresses limitations in current VLN systems by exp…

  8. RESEARCH · CL_193019 ·

    New benchmark VQABench analyzes cost-quality of cloud VLM VQA systems

    A new benchmark called VQABench has been developed to evaluate the cost-quality trade-offs of cloud-based Vision-Language Models (VLMs) for visual question answering (VQA) systems. The research highlights that client-si…

  9. RESEARCH · CL_187473 ·

    New frameworks and benchmarks advance video anomaly detection capabilities

    Researchers are developing advanced methods for video anomaly detection (VAD), a critical task for industrial applications and safety systems. New frameworks like VTO and FedVAR aim to improve generalization and address…

  10. TOOL · CL_185464 ·

    BIM-Native Tokenization Enhances Room Layout Synthesis

    Researchers have developed a novel BIM-native tokenization method for synthesizing room layouts within Building Information Modeling (BIM) scenes. This approach encodes each room as a sequence of BIM-Token Bundles, unif…

  11. TOOL · CL_185361 ·

    New research explores visual evidence representation for AI item difficulty prediction

    Researchers have explored methods for representing visual evidence in item difficulty prediction, comparing visual textualization (expressing images in language) with image-native modeling (retaining the original image)…

  12. RESEARCH · CL_186965 ·

    New GST-Bench benchmark reveals VLM struggles with global spatial awareness in video

    Researchers have introduced GST-Bench, a new benchmark designed to evaluate the global spatial awareness of Vision-Language Models (VLMs) using video data. The benchmark, which includes questions derived from over 6,790…

  13. TOOL · CL_183415 ·

    New CROSS method enhances remote sensing image segmentation

    Researchers have developed a new method called CROSS for referring remote sensing image segmentation. This approach aims to address limitations in existing Vision-Language Models (VLMs) and the Segment Anything Model (S…

  14. TOOL · CL_183403 ·

    New method enhances 3D indoor scene generation using graph validation

    Researchers have developed a new method called Global Graph-Validated Optimization for generating 3D indoor scenes from text instructions. This approach uses a graph-based representation to separate semantic coherence f…

  15. TOOL · CL_183293 ·

    New BanglaWild benchmark evaluates Bengali scene text recognition for OCR and VLMs

    Researchers have introduced BanglaWild, a new benchmark designed to evaluate Bengali scene text recognition for both optical character recognition (OCR) systems and vision-language models (VLMs). The benchmark consists …

  16. RESEARCH · CL_183045 ·

    New benchmarks and methods advance multimodal reasoning in AI

    Researchers are developing new methods for multimodal knowledge graph completion and reasoning, integrating vision-language models (VLMs) with graph structures. ViSR-KGC proposes a visual subgraph reasoning approach tha…

  17. RESEARCH · CL_183452 ·

    UniEvo-RS framework enhances remote sensing segmentation with exemplar-driven prototypes

    Researchers have developed UniEvo-RS, a novel framework for remote sensing segmentation that utilizes an omni-prompt approach with representative exemplar-driven prototype evolution. This system aims to overcome the per…

  18. RESEARCH · CL_183292 ·

    LLMs and VLMs show sensitivity to text casing, influencing attention

    A new research paper explores how Large Language Models (LLMs) and Vision-Language Models (VLMs) are sensitive to letter casing, similar to human visual perception. The study found that formatting text in uppercase or a…

  19. TOOL · CL_181139 ·

    New XSPA attack method targets Vision-Language Models

    Researchers have developed XSPA, a novel method for creating adversarial perturbations on Vision-Language Models (VLMs). This technique crafts imperceptible X-shaped sparse perturbations that can significantly degrade V…

  20. TOOL · CL_181060 ·

    New VC-Tooler framework enhances visual tool use for Vision--Language Models

    Researchers have introduced VC-Tooler, a new framework designed to enhance the capabilities of Vision--Language Models (VLMs) in utilizing visual tools. Unlike previous methods that focused on single-tool grounding, VC-…