PulseAugur
实时 11:02:39

新研究解决语言模型中的视听事件感知问题

两篇新研究论文介绍了语言模型中视听事件感知的新方法。第一篇ST-OmniQA提出了一个用于移动声源的时空视听推理基准,以及一个名为ST-Omni-R1的模型,该模型集成了视听上下文以改进事件识别和跟踪。第二篇SCoPE提供了一个无需训练的框架,利用稀疏的跨模态先验交换来增强视听事件感知,使模态能够相互引导,从而减轻虚假共激活。 AI

影响 这些进展可能导致更复杂的、能够更深入理解视频中动态事件的多模态AI系统。

排序理由 两篇在arXiv上发表的学术论文,介绍了用于视听事件感知的新基准和框架。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 3 个来源。 我们如何撰写摘要 →

新研究解决语言模型中的视听事件感知问题

报道来源 [3]

  1. arXiv cs.AI TIER_1 English(EN) · Zhi Zeng, Cheng Zhang, Zesheng Yang, Rendong Pi, Jiaying Wu, Di Zhang, Zihan Ma, Guodong Li, Zhou Yang, Yu Xiang, Yifei Zheng, Minnan Luo ·

    倾听、观察与追踪:全模态语言模型的时空视听声事件推理

    arXiv:2608.09435v1 Announce Type: new Abstract: Understanding dynamic sound sources requires jointly determining what produces a sound, where the source is located, and how it moves over time. Yet existing audio-language models often represent clips as global acoustic events, whi…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    倾听、观察与追踪:面向全模态语言模型的时空视听声事件推理

    Understanding dynamic sound sources requires jointly determining what produces a sound, where the source is located, and how it moves over time. Yet existing audio-language models often represent clips as global acoustic events, while vision-language models lack the spatial audio…

  3. arXiv cs.CV TIER_1 English(EN) · Jaemo Jeong, Junho Yoon, Hyunju Kim, Dongman Lee ·

    SCoPE:通过稀疏跨模态先验交换进行无需训练的视听事件感知

    arXiv:2608.07923v1 Announce Type: new Abstract: Audio-visual event perception (AVEP) determines which events occur in a video, when they occur, and whether they are audible, visible, or both. Training-free methods query new event vocabularies by matching frozen audio and visual f…