PulseAugur
中
实时 15:03:02
English(EN) CRAFT: Causal Responsibility and Failure Tracing in Medical Vision Language Models

视觉语言模型:规模、安全性和效率的探索

近期研究探讨了视觉语言模型(VLM)的复杂性,重点关注它们如何处理视觉和文本信息以及它们的局限性所在。研究调查了模型大小与视觉输入分辨率之间的权衡,医疗诊断应用中失败的因果机制,以及模态差距的几何特性。进一步的研究检查了VLM的注意力机制如何与人类注视对齐,并开发了高效的token剪枝方法以降低计算成本。此外,正在创建新的数据集和基准,以提高VLM在电路布局分析等专业领域的性能,并审计它们对字体图层的理解。 AI

影响 这些研究促进了对VLM能力、局限性和潜在失败模式的理解,为未来的开发和安全考量提供了信息。

排序理由 多篇arXiv论文提出了与视觉语言模型相关的新研究发现和方法论。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 17 个来源。 我们如何撰写摘要 →

视觉语言模型:规模、安全性和效率的探索

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇arXiv论文提出了与视觉语言模型相关的新研究发现和方法论。
Source corroboration
17 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, model release, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
6 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [17]

  1. arXiv cs.AI TIER_1 English(EN) · Xinye Zhao, Yunkai Dang, Yunchen Wu, Wenbin Li ·

    并非所有错误都能通过规模解决:视觉语言推理的规模极限在哪里

    arXiv:2610.01640v1 Announce Type: cross Abstract: Vision-language models (VLMs) face a fixed-budget trade-off between processing more visual information for fine-grained perception and using a larger language backbone for complex reasoning. Existing studies do not tell us which c…

  2. arXiv cs.AI TIER_1 English(EN) · Chunzheng Zhu, Jiaqi Zeng, Hongbo Zhao, Yihang Chen, Yijun Wang, Jianxin Lin ·

    CRAFT:医疗视觉语言模型中的因果责任与失败追踪

    arXiv:2609.38810v1 Announce Type: cross Abstract: As vision language models are increasingly deployed in clinical diagnosis, understanding how they internally resolve competing visual and textual signals becomes a safety imperative. Existing mechanistic analyses remain confined t…

  3. arXiv cs.AI TIER_1 English(EN) · Aditya Sharma, Divya Saxena ·

    一种几何,不同结果:视觉语言模型中模态差距的可读性依赖效应

    arXiv:2609.36101v1 Announce Type: cross Abstract: Contrastive vision-language models learn shared embedding spaces by aligning matched image-text pairs, yet their representations remain separated by a modality gap. Prior work reports divergent effects of modifying this gap: reduc…

  4. arXiv cs.AI TIER_1 English(EN) · Sumin Hong, Katsumi Ibaraki, Renee Shi, David Chiang, Toby Jia-Jun Li ·

    相似的选择,不同的注意力:人类与视觉语言模型的跨模态联想

    arXiv:2609.36475v1 Announce Type: cross Abstract: Cross-modal associations are systematic pairings of features across modalities, such as the association of 'bouba' with round shapes and 'kiki' with sharp shapes. Prior work has compared humans and vision-language models (VLMs) on…

  5. arXiv cs.AI TIER_1 English(EN) · Jing Wang, Zhiping Wu, Dongdong Ren, Youfang Han, Wei Zhao, Wenbin Li ·

    TReVS:整合文本相关性和视觉显著性以实现高效的视觉语言模型令牌剪枝

    arXiv:2609.37581v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) excel at visual understanding and reasoning but often incur substantial inference costs due to the large number of visual tokens. Recent visual token pruning methods increasingly follow a two-stage pa…

  6. arXiv cs.LG TIER_1 English(EN) · Woosang Jeon, Jiwon Yang, Soo Chung, Taehyeong Kim ·

    看见不等于解决:审计语言对冻结视觉几何的访问

    arXiv:2609.37230v1 Announce Type: cross Abstract: Visual distinctions are often finer than those reflected in linguistic conceptualization. Vision-language models exhibit a similar asymmetry: a distinction can remain discriminable in frozen image geometry while being weakly addre…

  7. arXiv cs.LG TIER_1 English(EN) · Giuseppe Chiari, Michele Piccoli, Federico Viola, Davide Zoni ·

    THEIA:用于视觉语言布局分析的多模态数据集和基准测试

    arXiv:2609.35035v2 Announce Type: replace Abstract: The integration of artificial intelligence into computer-aided design frameworks has sparked a shift in the design of analog integrated circuits (ICs), transitioning the field from using manual and algorithmic-based solutions to…

  8. Hugging Face Daily Papers TIER_1 English(EN) ·

    CRAFT:医疗视觉语言模型中的因果责任与失败追踪

    As vision language models are increasingly deployed in clinical diagnosis, understanding how they internally resolve competing visual and textual signals becomes a safety imperative. Existing mechanistic analyses remain confined to unimodal text and offer no explanation for why a…

  9. arXiv cs.AI TIER_1 English(EN) · Mert \.Incidelen, Yamen Kashkash, Asya Berker, Murat Aydo\u{g}an ·

    抱歉,机器人,快乐的人类:视觉语言模型只能读取两个可读的排版层中的一个

    arXiv:2609.31403v1 Announce Type: cross Abstract: Vision-language models (VLMs), despite their success in optical character recognition (OCR) tasks, are vulnerable to typographic attacks and have a fragile structure for images with multiple text layers. In this study, the DecoyBe…

  10. Hugging Face Daily Papers TIER_1 English(EN) ·

    当文本至关重要时:视觉语言模型中视觉标记修剪的设计原则

    Visual token pruning has been widely studied as a practical approach to reducing the computational cost of large vision-language models. However, it struggles to preserve essential visual information, which can lead to substantial performance degradation. In particular, image-bas…

  11. arXiv cs.CV TIER_1 English(EN) · Genpei Zhang ·

    编码但断开:在打补丁空值下分解视觉语言模型的失败

    arXiv:2610.00024v1 Announce Type: new Abstract: Across three vision-language model architectures (LLaVA-1.5-7B, Qwen2.5-VL-7B, InternVL3-8B), we report a universal negative finding for mid-layer interpretability. On POPE -- the benchmark common to all three -- the mid layers enco…

  12. arXiv cs.CV TIER_1 English(EN) · Elias Rotondo (Duke University), Lin Duan (Duke University), Yanming Xiu (Duke University), Sangjun Eom (Duke University), Conrad Li (Duke University), Maria Gorlatova (Duke University) ·

    利用视觉语言模型进行增强现实中的感知质量评估和自主内容调整

    arXiv:2610.00677v1 Announce Type: new Abstract: Advancements in augmented reality (AR) continue to foster innovative solutions, facilitating novel methodologies within educational systems, healthcare delivery, and risk-mitigation protocols. However, optimizing for end-user immers…

  13. arXiv cs.CV TIER_1 (CA) · Guo Cheng, Huang Li ·

    感知退化下视觉-语言模型的语义失调

    arXiv:2601.08355v3 Announce Type: replace Abstract: Vision-Language Models (VLMs) are increasingly deployed in autonomous driving and embodied AI systems, where reliable perception is critical for safe semantic reasoning and decision-making. While recent VLMs demonstrate strong p…

  14. arXiv cs.CV TIER_1 English(EN) · Binchi Zhang, Atrisha Sarkar, Apurva Narayan ·

    只需微小改动即可重塑认知:在 $\epsilon \leq 4/255$ 下对视觉语言模型进行目标语义替换

    arXiv:2609.38298v1 Announce Type: new Abstract: Vision Language Models (VLMs) are widely deployed in safety-critical scenarios, and understanding to which extent they can be controlled by adversarial perturbation is a prerequisite for evaluating their trustworthiness. Existing re…

  15. arXiv cs.CV TIER_1 English(EN) · Weixuan Li, Zikun Zhou, Xinyi Zhuang, Xinyan Guo, Rui Tian, Chuyao Zhang, Lin Gao ·

    表示动力学揭示多模态大模型视觉Token剪枝的语义显著性和相似性

    arXiv:2609.36916v1 Announce Type: new Abstract: Multimodal large language models (MLLMs) incur high inference latency from long visual token sequences. Existing pruning methods commonly use attention maps or output features to estimate token importance or redundancy. Several rece…

  16. arXiv cs.CV TIER_1 English(EN) · Van Bach Nguyen, J\"org Schl\"otterer, Christin Seifer ·

    面向对比视觉语言模型的定向视觉反事实解释

    arXiv:2609.37638v1 Announce Type: new Abstract: Current explanation methods for contrastive vision--language models such as CLIP mainly identify important regions without showing how to change the input in order to get a target prediction. We introduce \textbf{M}ask-guided \textb…

  17. arXiv cs.CV TIER_1 English(EN) · Bowen Yuan, Danny Wang, Ruihong Qiu, Zijian Wang, Zi Huang ·

    追踪证据:通过视觉语言推理实现忠实的 Token 归因

    arXiv:2609.37656v1 Announce Type: new Abstract: Large vision-language models (LVLMs) exhibit strong reasoning capabilities, yet the visual and textual evidence supporting the generated responses remains difficult to identify. Faithful token attribution explains an LVLM's response…