PulseAugur
实时 02:45:40
English(EN) Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models

新方法提升全双工语音模型以实现更好的交互

研究人员开发了新的方法来增强全双工语音模型,从而实现更自然、更具交互性的对话。一种方法侧重于使用强化学习来改进暂停处理和轮流发言等交互轴,并将其应用于 MoshiPersonaPlex 等模型。另一种方法,Listen-Write-Speak (LWS),引入了一种以文本为中心的范式,模型可以同时收听、写入可见文本并进行语音输出,利用文本原生能力而不牺牲实时响应能力。 AI

影响 这些进步可能带来更自然、更强大的语音助手和对话式AI系统。

排序理由 该集群包含两篇研究论文,详细介绍了全双工语音模型的新方法。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 5 个来源。 我们如何撰写摘要 →

新方法提升全双工语音模型以实现更好的交互

报道来源 [5]

  1. arXiv cs.AI TIER_1 English(EN) · Cheng-Kuang Chang, Kai-Wei Chang, Alexander H. Liu, James Glass ·

    通过激活引导克服全双工口语语言模型中的状态惯性

    arXiv:2606.11386v1 Announce Type: cross Abstract: Full-duplex spoken language models (FD-SLMs) enable seamless speech interaction by allowing models to listen and speak simultaneously, yet the internal mechanism by which they coordinate listening and speaking remains underexplore…

  2. arXiv cs.CL TIER_1 English(EN) · Atsumoto Ohashi, Neil Zeghidour, Alexandre D\'efossez, Eugene Kharitonov ·

    全双工语音模型中的多面交互对齐

    arXiv:2606.11167v1 Announce Type: new Abstract: Full-duplex spoken dialogue models can listen and speak simultaneously, making them a promising architecture for natural conversation. However, current models are trained solely with supervised learning through token-level likelihoo…

  3. arXiv cs.CL TIER_1 English(EN) · Eugene Kharitonov ·

    全双工语音模型中的多面交互对齐

    Full-duplex spoken dialogue models can listen and speak simultaneously, making them a promising architecture for natural conversation. However, current models are trained solely with supervised learning through token-level likelihood maximization, which does not directly optimize…

  4. arXiv cs.AI TIER_1 English(EN) · Luoyuan Zhang, Bokai Xu, Junbo Cui, Weiyue Sun, Yingjing Xu, Hanyu Liu, Yuan Yao ·

    在全双工语音模型中释放 LLM 的能力

    arXiv:2606.07547v1 Announce Type: cross Abstract: Speech-based large language models are typically constrained to spoken replies, which limits their user-facing outputs to what can be verbalized and suppresses text-native capabilities such as code generation, structured analysis,…

  5. Hugging Face Daily Papers TIER_1 English(EN) ·

    在全双工语音模型中释放 LLM 的能力

    A text-first tri-channel speech interface enables real-time interaction with visible text output alongside spoken responses, demonstrating superior performance in full-duplex conversational tasks.