Two new research papers introduce novel approaches to audio-visual event perception in language models. The first, ST-OmniQA, presents a benchmark for spatio-temporal audio-visual reasoning with moving sound sources, alongside a model, ST-Omni-R1, that integrates audio-visual context for improved event recognition and tracking. The second, SCoPE, offers a training-free framework that leverages sparse cross-modal prior exchange to enhance audio-visual event perception by allowing modalities to guide each other, thereby mitigating false co-activations. AI
IMPACT These advancements could lead to more sophisticated multi-modal AI systems capable of deeper understanding of dynamic events in videos.
RANK_REASON Two academic papers published on arXiv introducing new benchmarks and frameworks for audio-visual event perception.
Read on Hugging Face Daily Papers →
- alphaXiv
- arXiv
- AV$^2$A
- CatalyzeX
- CLAP
- Connected Papers
- DagsHub
- Gotit.pub
- Hugging Face
- Litmaps
- OV-AVEBench
- ScienceCast
- Scite
- SCoPE
- ST-OmniQA
- ST-Omni-R1
- VGGSound-AVEL100k
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →