English(EN)LOCUS: Local Visual Cue Search for Enhancing Fine-Grained Perception in Multimodal Large Language Models
PerceptionDLM实现多模态模型的并行区域感知 · 跟踪6个来源
作者PulseAugur 编辑部·[6 个来源]·
研究人员推出了PerceptionDLM,这是一种多模态扩散语言模型,专为高效的并行区域感知而设计。该模型利用扩散语言模型的并行解码能力,同时为多个图像区域生成描述,与顺序方法相比,显著提高了推理效率。为了评估这一新功能,创建了一个并行详细局部字幕基准(ParaDLC-Bench),该基准扩展了现有基准,使其包含每张图像的多个区域掩码。实验表明,PerceptionDLM在保持具有竞争力的字幕性能的同时,为多区域感知任务提供了显著的速度提升。
AI
arXiv:2606.19534v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks. However, most existing MLLMs rely on autoregressive generation, which limits their efficiency for perception tasks that requ…
Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks. However, most existing MLLMs rely on autoregressive generation, which limits their efficiency for perception tasks that require captioning multiple regions. In this work, we …
PerceptionDLM enables efficient parallel region perception in multimodal diffusion language models through structured attention masking and efficient prompting, achieving faster inference without sacrificing caption quality.
arXiv:2412.21036v3 Announce Type: replace Abstract: Geometric shapes play important roles in both physical world and human cognition. While multimodal large language models (MLLMs) have made significant advancements in visual understanding, their abilities to recognize geometric …
arXiv:2606.16586v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) remain unreliable on fine-grained visual perception, even when high-resolution inputs preserve the necessary local details. We identify this limitation as visual context rot: decisive evidenc…
Multimodal Large Language Models (MLLMs) remain unreliable on fine-grained visual perception, even when high-resolution inputs preserve the necessary local details. We identify this limitation as visual context rot: decisive evidence may exist in the full image, yet fail to be re…