PulseAugur
实时 19:11:54
English(EN) LOCUS: Local Visual Cue Search for Enhancing Fine-Grained Perception in Multimodal Large Language Models

PerceptionDLM实现多模态模型的并行区域感知 · 跟踪6个来源

研究人员推出了PerceptionDLM,这是一种多模态扩散语言模型,专为高效的并行区域感知而设计。该模型利用扩散语言模型的并行解码能力,同时为多个图像区域生成描述,与顺序方法相比,显著提高了推理效率。为了评估这一新功能,创建了一个并行详细局部字幕基准(ParaDLC-Bench),该基准扩展了现有基准,使其包含每张图像的多个区域掩码。实验表明,PerceptionDLM在保持具有竞争力的字幕性能的同时,为多区域感知任务提供了显著的速度提升。 AI

影响 这项研究可能带来更高效的多模态人工智能系统,实现更快、更详细的视觉理解。

排序理由 该集群包含多篇关于多模态大语言模型新模型和基准的研究论文。

在 arXiv cs.CV 阅读 →

AI 生成摘要 · Google Gemini · 来自 6 个来源。 我们如何撰写摘要 →

PerceptionDLM实现多模态模型的并行区域感知 · 跟踪6个来源

报道来源 [6]

  1. arXiv cs.AI TIER_1 English(EN) · Yueyi Sun, Yuhao Wang, Jason Li, Ye Tian, Tao Zhang, Jacky Mai, Yihan Wang, Haochen Wang, Jinbin Bai, Ling Yang, Yunhai Tong ·

    PerceptionDLM:多模态扩散语言模型的并行区域感知

    arXiv:2606.19534v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks. However, most existing MLLMs rely on autoregressive generation, which limits their efficiency for perception tasks that requ…

  2. arXiv cs.CL TIER_1 English(EN) · Yunhai Tong ·

    PerceptionDLM:多模态扩散语言模型的并行区域感知

    Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks. However, most existing MLLMs rely on autoregressive generation, which limits their efficiency for perception tasks that require captioning multiple regions. In this work, we …

  3. Hugging Face Daily Papers TIER_1 English(EN) ·

    PerceptionDLM:多模态扩散语言模型实现并行区域感知

    PerceptionDLM enables efficient parallel region perception in multimodal diffusion language models through structured attention masking and efficient prompting, achieving faster inference without sacrificing caption quality.

  4. arXiv cs.CL TIER_1 English(EN) · Shangyu Xing, Changhao Xiang, Yuteng Han, Yifan Yue, Zhen Wu, Xinyu Liu, Zhangtai Wu, Fei Zhao, Xinyu Dai ·

    GePBench:评估多模态大语言模型的基础几何感知能力

    arXiv:2412.21036v3 Announce Type: replace Abstract: Geometric shapes play important roles in both physical world and human cognition. While multimodal large language models (MLLMs) have made significant advancements in visual understanding, their abilities to recognize geometric …

  5. arXiv cs.CV TIER_1 English(EN) · Zhou Tao, Fang Zhang, Zewen Ding, Shida Wang, Xiaokun Sun, YongXiang Hua, Haoyu Cao, Linli Xu ·

    LOCUS:用于增强多模态大语言模型细粒度感知的局部视觉线索搜索

    arXiv:2606.16586v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) remain unreliable on fine-grained visual perception, even when high-resolution inputs preserve the necessary local details. We identify this limitation as visual context rot: decisive evidenc…

  6. arXiv cs.CV TIER_1 English(EN) · Linli Xu ·

    LOCUS:用于增强多模态大语言模型细粒度感知的局部视觉线索搜索

    Multimodal Large Language Models (MLLMs) remain unreliable on fine-grained visual perception, even when high-resolution inputs preserve the necessary local details. We identify this limitation as visual context rot: decisive evidence may exist in the full image, yet fail to be re…