Vision-Language Models: Scaling, Safety, and Efficiency Explored
ByPulseAugur Editorial·[17 sources]·
Recent research explores the intricacies of vision-language models (VLMs), focusing on how they process visual and textual information and where their limitations lie. Studies investigate the trade-offs between model size and visual input resolution, the causal mechanisms behind failures in medical diagnosis applications, and the geometric properties of the modality gap. Further research examines how VLMs' attention mechanisms align with human gaze and develops methods for efficient token pruning to reduce computational costs. Additionally, new datasets and benchmarks are being created to improve VLM performance in specialized domains like circuit layout analysis and to audit their understanding of typographic layers.
AI
IMPACT
These studies advance understanding of VLM capabilities, limitations, and potential failure modes, informing future development and safety considerations.
RANK_REASON
Multiple arXiv papers present novel research findings and methodologies related to vision-language models.
arXiv:2610.01640v1 Announce Type: cross Abstract: Vision-language models (VLMs) face a fixed-budget trade-off between processing more visual information for fine-grained perception and using a larger language backbone for complex reasoning. Existing studies do not tell us which c…
arXiv:2609.38810v1 Announce Type: cross Abstract: As vision language models are increasingly deployed in clinical diagnosis, understanding how they internally resolve competing visual and textual signals becomes a safety imperative. Existing mechanistic analyses remain confined t…
arXiv:2609.36475v1 Announce Type: cross Abstract: Cross-modal associations are systematic pairings of features across modalities, such as the association of 'bouba' with round shapes and 'kiki' with sharp shapes. Prior work has compared humans and vision-language models (VLMs) on…
arXiv:2609.37581v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) excel at visual understanding and reasoning but often incur substantial inference costs due to the large number of visual tokens. Recent visual token pruning methods increasingly follow a two-stage pa…
arXiv cs.LG
TIER_1English(EN)·Woosang Jeon, Jiwon Yang, Soo Chung, Taehyeong Kim·
arXiv:2609.37230v1 Announce Type: cross Abstract: Visual distinctions are often finer than those reflected in linguistic conceptualization. Vision-language models exhibit a similar asymmetry: a distinction can remain discriminable in frozen image geometry while being weakly addre…
arXiv:2609.35035v2 Announce Type: replace Abstract: The integration of artificial intelligence into computer-aided design frameworks has sparked a shift in the design of analog integrated circuits (ICs), transitioning the field from using manual and algorithmic-based solutions to…
As vision language models are increasingly deployed in clinical diagnosis, understanding how they internally resolve competing visual and textual signals becomes a safety imperative. Existing mechanistic analyses remain confined to unimodal text and offer no explanation for why a…
arXiv cs.AI
TIER_1English(EN)·Mert \.Incidelen, Yamen Kashkash, Asya Berker, Murat Aydo\u{g}an·
arXiv:2609.31403v1 Announce Type: cross Abstract: Vision-language models (VLMs), despite their success in optical character recognition (OCR) tasks, are vulnerable to typographic attacks and have a fragile structure for images with multiple text layers. In this study, the DecoyBe…
Visual token pruning has been widely studied as a practical approach to reducing the computational cost of large vision-language models. However, it struggles to preserve essential visual information, which can lead to substantial performance degradation. In particular, image-bas…
arXiv:2610.00024v1 Announce Type: new Abstract: Across three vision-language model architectures (LLaVA-1.5-7B, Qwen2.5-VL-7B, InternVL3-8B), we report a universal negative finding for mid-layer interpretability. On POPE -- the benchmark common to all three -- the mid layers enco…
arXiv cs.CV
TIER_1English(EN)·Elias Rotondo (Duke University), Lin Duan (Duke University), Yanming Xiu (Duke University), Sangjun Eom (Duke University), Conrad Li (Duke University), Maria Gorlatova (Duke University)·
arXiv:2610.00677v1 Announce Type: new Abstract: Advancements in augmented reality (AR) continue to foster innovative solutions, facilitating novel methodologies within educational systems, healthcare delivery, and risk-mitigation protocols. However, optimizing for end-user immers…
arXiv:2601.08355v3 Announce Type: replace Abstract: Vision-Language Models (VLMs) are increasingly deployed in autonomous driving and embodied AI systems, where reliable perception is critical for safe semantic reasoning and decision-making. While recent VLMs demonstrate strong p…
arXiv:2609.38298v1 Announce Type: new Abstract: Vision Language Models (VLMs) are widely deployed in safety-critical scenarios, and understanding to which extent they can be controlled by adversarial perturbation is a prerequisite for evaluating their trustworthiness. Existing re…
arXiv:2609.36916v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) incur high inference latency from long visual token sequences. Existing pruning methods commonly use attention maps or output features to estimate token importance or redundancy. Several rece…
arXiv:2609.37638v1 Announce Type: new Abstract: Current explanation methods for contrastive vision--language models such as CLIP mainly identify important regions without showing how to change the input in order to get a target prediction. We introduce \textbf{M}ask-guided \textb…
arXiv cs.CV
TIER_1English(EN)·Bowen Yuan, Danny Wang, Ruihong Qiu, Zijian Wang, Zi Huang·
arXiv:2609.37656v1 Announce Type: new Abstract: Large vision-language models (LVLMs) exhibit strong reasoning capabilities, yet the visual and textual evidence supporting the generated responses remains difficult to identify. Faithful token attribution explains an LVLM's response…