PulseAugur
EN
LIVE 08:21:07

New research tackles audio-visual event perception in language models

Two new research papers introduce novel approaches to audio-visual event perception in language models. The first, ST-OmniQA, presents a benchmark for spatio-temporal audio-visual reasoning with moving sound sources, alongside a model, ST-Omni-R1, that integrates audio-visual context for improved event recognition and tracking. The second, SCoPE, offers a training-free framework that leverages sparse cross-modal prior exchange to enhance audio-visual event perception by allowing modalities to guide each other, thereby mitigating false co-activations. AI

IMPACT These advancements could lead to more sophisticated multi-modal AI systems capable of deeper understanding of dynamic events in videos.

RANK_REASON Two academic papers published on arXiv introducing new benchmarks and frameworks for audio-visual event perception.

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

New research tackles audio-visual event perception in language models

COVERAGE [2]

  1. arXiv cs.AI TIER_1 English(EN) · Zhi Zeng, Cheng Zhang, Zesheng Yang, Rendong Pi, Jiaying Wu, Di Zhang, Zihan Ma, Guodong Li, Zhou Yang, Yu Xiang, Yifei Zheng, Minnan Luo ·

    Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models

    arXiv:2608.09435v1 Announce Type: new Abstract: Understanding dynamic sound sources requires jointly determining what produces a sound, where the source is located, and how it moves over time. Yet existing audio-language models often represent clips as global acoustic events, whi…

  2. arXiv cs.CV TIER_1 English(EN) · Jaemo Jeong, Junho Yoon, Hyunju Kim, Dongman Lee ·

    SCoPE: Training-Free Audio-Visual Event Perception via Sparse Cross-Modal Prior Exchange

    arXiv:2608.07923v1 Announce Type: new Abstract: Audio-visual event perception (AVEP) determines which events occur in a video, when they occur, and whether they are audible, visible, or both. Training-free methods query new event vocabularies by matching frozen audio and visual f…