PulseAugur
实时 19:24:56
English(EN) MUGEN: Evaluating and Improving Multi-audio Understanding of Large Audio-Language Models

新研究解决音频语言模型在指令遵循和评估方面的局限性

研究人员正在开发新方法来改进大型音频语言模型(LALM)的功能。一种方法侧重于使用音频感知LLM为文本到音频生成中的指令遵循提供细粒度反馈,特别是在涉及多个声音事件和时间顺序的任务中。另一项研究调查了LALM在语音评估中可能采取的潜在捷径,发现一些模型依赖于协议级别的线索,而不是真正地收听音频。此外,还引入了新的基准和模型来评估和增强长音频录音中的多音频理解和时间定位能力。 AI

影响 音频语言模型的进步可能导致更复杂的AI系统能够理解和生成复杂的音频,从而影响从内容创建到可访问性的各种应用。

排序理由 arXiv上发表了多篇研究论文,详细介绍了大型音频语言模型的新方法、基准和评估。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 5 个来源。 我们如何撰写摘要 →

新研究解决音频语言模型在指令遵循和评估方面的局限性

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
arXiv上发表了多篇研究论文,详细介绍了大型音频语言模型的新方法、基准和评估。
Source corroboration
5 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
53 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准

报道来源 [5]

  1. arXiv cs.AI TIER_1 English(EN) · Chun-Yi Kuan, Siwon Kim, Byeonggeun Kim, Suyoun Kim, Bo-Ru Lu, Qinming Tang, Ankur Gandhe, Hung-yi Lee, Chieh-Chi Kao, Chao Wang ·

    通过音频感知大语言模型的细粒度反馈改进文本到音频指令遵循

    arXiv:2607.13408v1 Announce Type: cross Abstract: Recent text-to-audio models generate high-quality audio, but often fail to follow instructions involving multiple sound events and temporal order. This gap arises because existing evaluation and training signals mainly emphasize g…

  2. arXiv cs.CL TIER_1 English(EN) · Hiroshi Saruwatari ·

    审计大型音频语言模型评判器中的协议级捷径以进行语音评估

    Large audio-language models (LALMs) are increasingly used as automatic judges for speech evaluation. However, high agreement with human ratings does not guarantee that their verdicts are grounded in the audio. A judge may instead rely on specialist labels or reference data suppli…

  3. arXiv cs.AI TIER_1 English(EN) · Chao Wang ·

    通过音频感知大语言模型的细粒度反馈改进文本到音频指令遵循

    Recent text-to-audio models generate high-quality audio, but often fail to follow instructions involving multiple sound events and temporal order. This gap arises because existing evaluation and training signals mainly emphasize global similarity or perceptual quality, with limit…

  4. arXiv cs.AI TIER_1 English(EN) · Chih-Kai Yang, Yun-Shao Tsai, Yu-Kai Guo, Ping-Le Tsai, Yen-Ting Piao, Hung-Wei Chen, Ting-Lin Hsiao, Yun-Man Hsu, Ke-Han Lu, Hung-yi Lee ·

    MUGEN:评估和改进大型音频语言模型的多音频理解能力

    arXiv:2603.09714v2 Announce Type: replace-cross Abstract: While multi-audio understanding is critical for large audio-language models (LALMs), it remains underexplored. We introduce MUGEN, a comprehensive benchmark evaluating this capability across speech, general audio, and musi…

  5. arXiv cs.CL TIER_1 English(EN) · Aleksandr Kutsakov, Mariia Sadovina, Georgii Gospodinov, Alexandr Maximenko, Oleg Kutuzov, Pavel Bogomolov, Fyodor Minkin ·

    GigaChat Audio:面向时间的超大音频语言模型

    arXiv:2607.10387v1 Announce Type: cross Abstract: Temporal grounding in long recordings remains challenging for audio-conditioned LLMs. We present a time-aware audio LLM that answers questions with explicit timestamps over up to 120 minutes of input. Our approach interleaves peri…