PulseAugur
EN
LIVE 03:47:29

New methods boost full-duplex speech models for better interaction

Researchers have developed new methods to enhance full-duplex speech models, enabling more natural and interactive conversations. One approach focuses on improving interactivity axes like pause handling and turn-taking using reinforcement learning, applied to models like Moshi and PersonaPlex. Another method, Listen-Write-Speak (LWS), introduces a text-first paradigm where models can simultaneously listen, write visible text, and speak, leveraging text-native capabilities without sacrificing real-time responsiveness. AI

IMPACT These advancements could lead to more natural and capable voice assistants and conversational AI systems.

RANK_REASON The cluster contains two research papers detailing new methods for full-duplex speech models.

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 5 sources. How we write summaries →

New methods boost full-duplex speech models for better interaction

COVERAGE [5]

  1. arXiv cs.AI TIER_1 English(EN) · Cheng-Kuang Chang, Kai-Wei Chang, Alexander H. Liu, James Glass ·

    Overcoming State Inertia in Full-Duplex Spoken Language Models via Activation Steering

    arXiv:2606.11386v1 Announce Type: cross Abstract: Full-duplex spoken language models (FD-SLMs) enable seamless speech interaction by allowing models to listen and speak simultaneously, yet the internal mechanism by which they coordinate listening and speaking remains underexplore…

  2. arXiv cs.CL TIER_1 English(EN) · Atsumoto Ohashi, Neil Zeghidour, Alexandre D\'efossez, Eugene Kharitonov ·

    Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models

    arXiv:2606.11167v1 Announce Type: new Abstract: Full-duplex spoken dialogue models can listen and speak simultaneously, making them a promising architecture for natural conversation. However, current models are trained solely with supervised learning through token-level likelihoo…

  3. arXiv cs.CL TIER_1 English(EN) · Eugene Kharitonov ·

    Multi-Faceted Interactivity Alignment in Full-Duplex Speech Models

    Full-duplex spoken dialogue models can listen and speak simultaneously, making them a promising architecture for natural conversation. However, current models are trained solely with supervised learning through token-level likelihood maximization, which does not directly optimize…

  4. arXiv cs.AI TIER_1 English(EN) · Luoyuan Zhang, Bokai Xu, Junbo Cui, Weiyue Sun, Yingjing Xu, Hanyu Liu, Yuan Yao ·

    Liberating LLM Capabilities in Full-Duplex Speech Models

    arXiv:2606.07547v1 Announce Type: cross Abstract: Speech-based large language models are typically constrained to spoken replies, which limits their user-facing outputs to what can be verbalized and suppresses text-native capabilities such as code generation, structured analysis,…

  5. Hugging Face Daily Papers TIER_1 English(EN) ·

    Liberating LLM Capabilities in Full-Duplex Speech Models

    A text-first tri-channel speech interface enables real-time interaction with visible text output alongside spoken responses, demonstrating superior performance in full-duplex conversational tasks.