English(EN)MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models
新研究解决音频语言模型在指令遵循和评估方面的局限性
作者PulseAugur 编辑部·[5 个来源]·
研究人员正在开发新方法来改进大型音频语言模型(LALM)的功能。一种方法侧重于使用音频感知LLM为文本到音频生成中的指令遵循提供细粒度反馈,特别是在涉及多个声音事件和时间顺序的任务中。另一项研究调查了LALM在语音评估中可能采取的潜在捷径,发现一些模型依赖于协议级别的线索,而不是真正地收听音频。此外,还引入了新的基准和模型来评估和增强长音频录音中的多音频理解和时间定位能力。
AI
arXiv:2607.13408v1 Announce Type: cross Abstract: Recent text-to-audio models generate high-quality audio, but often fail to follow instructions involving multiple sound events and temporal order. This gap arises because existing evaluation and training signals mainly emphasize g…
Large audio-language models (LALMs) are increasingly used as automatic judges for speech evaluation. However, high agreement with human ratings does not guarantee that their verdicts are grounded in the audio. A judge may instead rely on specialist labels or reference data suppli…
Recent text-to-audio models generate high-quality audio, but often fail to follow instructions involving multiple sound events and temporal order. This gap arises because existing evaluation and training signals mainly emphasize global similarity or perceptual quality, with limit…
arXiv:2603.09714v2 Announce Type: replace-cross Abstract: While multi-audio understanding is critical for large audio-language models (LALMs), it remains underexplored. We introduce MUGEN, a comprehensive benchmark evaluating this capability across speech, general audio, and musi…
arXiv cs.CL
TIER_1English(EN)·Aleksandr Kutsakov, Mariia Sadovina, Georgii Gospodinov, Alexandr Maximenko, Oleg Kutuzov, Pavel Bogomolov, Fyodor Minkin·
arXiv:2607.10387v1 Announce Type: cross Abstract: Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input. Our approach interleaves peri…