PulseAugur
中
实时 12:39:26
English(EN) Enabling Streaming User Transcription in Full-Duplex Speech-to-Speech Models

新数据集和模型推动全双工口语对话研究

研究人员推出了专注于全双工口语对话系统的新数据集和模型,这些系统能够实现更自然、实时的对话交互。DuplexDrama 数据集提供了超过 2000 小时的合成音频数据,涵盖场景、富有表现力的语音和声音事件,并将发布部分双语对话。SteerDuplex 引入了一个用于可控全双工语音的模型和基准,允许控制音调和语速等属性,并在指令遵循和中断处理方面取得了显著改进。此外,ConversationalVoice 提供了一个从真实对话创建训练数据的流水线,生成保留交互动态的分离、重建和扩展语音片段。 AI

影响 全双工口语对话系统的这些进步可能带来更自然、响应更快的 AI 助手和对话代理。

排序理由 多篇研究论文介绍了口语对话系统的新数据集和模型。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 5 个来源。 我们如何撰写摘要 →

新数据集和模型推动全双工口语对话研究

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
多篇研究论文介绍了口语对话系统的新数据集和模型。
Source corroboration
5 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
paper, model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
30 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.
Coverage growth since scoring
+2 source(s) since last score
New sources have picked up this story since our last re-score. Score will update on the next scoring pass.

完整方法见我们的编辑标准。

报道来源 [5]

  1. arXiv cs.CL TIER_1 English(EN) · Shuofeng Zhao, Hongwei Cai, Wenke Fan, Qingxiang Guo, Dawei Yang, Zhou Wang, Zhiyang Zhou, Yingxin Shang, Weixu Wang, Lin Yang, Shuran Zhou, Yang Song ·

    ECHO: 一个用于全双工对话中上下文敏感轮次切换的匹配对比基准

    arXiv:2609.17360v1 Announce Type: new Abstract: Full-duplex spoken dialogue systems must distinguish interruptions that require yielding the floor from backchannels that permit continued speaking. Existing benchmarks typically evaluate events independently and may therefore rewar…

  2. arXiv cs.CL TIER_1 English(EN) · Ke Hu, Nourchene Ferchichi, Edresson Casanova, Ankita Pasad, Elena Rastorgueva, Chen Chen, Nithin Rao Koluguri, Piotr Zelasko, Yifan Peng, Hainan Xu, Zhehuai Chen, Boris Ginsburg ·

    赋能全双工语音到语音模型中的流式用户转录

    arXiv:2609.15759v1 Announce Type: new Abstract: Full-duplex speech-to-speech (S2S) models enable natural conversational AI by allowing simultaneous listening and speaking. However, these models typically lack inherent user speech transcription, which is essential for applications…

  3. arXiv cs.CL TIER_1 English(EN) · Qingxiang Guo, Wenke Fan, Shuofeng Zhao, Dawei Yang, Zhiyang Zhou, Yingxin Shang, Hongwei Cai, Zhou Wang, Weixu Wang, Lin Yang, Shuran Zhou, Yang Song ·

    DuplexDrama:一个包含场景、全双工行为、富有表现力的语音和声音事件的合成对话数据集

    arXiv:2609.12872v1 Announce Type: new Abstract: We present DuplexDrama, the first synthesized spoken dialogue dataset that simultaneously covers four dimensions: (i) complete persona and scenario settings; (ii) three full-duplex behaviors (interruption, backchannel, incomplete); …

  4. arXiv cs.CL TIER_1 English(EN) · Utkarsh Tyagi, Ramaneswaran Selvakumar, Advait Gosai, Sonal Kumar, Nikhil Barhate, Isabell Sagar, Steven Li, Miheer Bavare, Daniel Quigley, Fabiola Tapia Carrillo, Jose M Patron E, Diego Mac\'ias Guti\'errez, Paul Song, Ramani Duraiswami, Dinesh Manocha,… ·

    SteerDuplex:可控双工语音对话模型

    arXiv:2609.12623v1 Announce Type: cross Abstract: Full-duplex spoken dialogue models support low-latency turn taking, interruption handling, and backchanneling, yet a key capability remains underexplored: steerability, the ability to reliably shift conversational behavior along a…

  5. Hugging Face Daily Papers TIER_1 English(EN) ·

    ConversationalVoice:通过源忠实重建和对话式扩展实现的真实对话全双工语音数据

    Full-duplex speech models require training data that preserves turn-taking, overlap, interruption, and backchannel behavior, yet these signals are entangled across speakers in noisy real-world recordings. We present Conversational Voice, a pipeline that converts real two-speaker …