PulseAugur
中
实时 23:33:43
English(EN) Audio LLMs Know When They Can't Hear You

音频大语言模型现可检测不可靠的转录

研究人员开发了一种方法,使音频大语言模型(LLM)能够识别它们何时无法可靠地转录用户输入。当前方法,包括语音质量预测器和生成不确定性,效果有限。然而,研究发现,LLM中的音频编码器表示强烈指示了转录的可靠性。一种新的、轻量级的预测器使用这些表示来识别不可靠的查询,从而使模型能够在生成错误响应之前向用户请求澄清。该预测器达到了很高的宏F1分数,并展示了跨不同音频LLM家族的可迁移性。 AI

影响 通过使模型能够在输入不清楚时请求澄清,增强了用户体验和语音AI交互的可靠性。

排序理由 详细介绍音频大语言模型新方法的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

音频大语言模型现可检测不可靠的转录

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Amirhosein Javadi, Richa Dixit, Mehrdad Farajtabar, Minsik Cho, Devang Naik, Mohammad Samragh ·

    音频大模型知道何时听不清你说话

    arXiv:2609.30625v1 Announce Type: new Abstract: Audio large language models allow users to interact with the model through speech. When an input recording is too degraded, the model may misinterpret the user's query and respond based on an incorrect transcription. In this paper, …