PulseAugur
实时 12:14:09

新的TTS研究探索离散流匹配以提高效率

两篇新研究论文探讨了零样本文本到语音(TTS)技术的进展,重点关注离散流匹配技术。第一篇论文介绍了DiFlow-TTS,一个使用离散流匹配方法来平衡生成质量和推理效率的框架,解决了自回归和连续空间基于流的模型存在的局限性。第二篇论文《Mask, Sample, Revise》提出了一种用于离散流匹配TTS的推理时堆栈,在不显式使用持续时间预测器的情况下,增强了从神经编解码器令牌生成语音的控制力和鲁棒性。 AI

影响 这些论文引入了文本到语音合成的新技术,可能带来更高效、更高质量的语音生成系统。

排序理由 两篇在arXiv上发表的学术论文,详细介绍了文本到语音合成的新方法。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

新的TTS研究探索离散流匹配以提高效率

报道来源 [2]

  1. arXiv cs.CL TIER_1 English(EN) · Ngoc-Son Nguyen, Thanh V. T. Tran, Hieu-Nghia Huynh-Nguyen, Truong-Son Hy, Van Nguyen ·

    DiFlow-TTS:具有离散流匹配的紧凑低延迟零样本文本到语音

    arXiv:2509.09631v4 Announce Type: replace-cross Abstract: Zero-shot text-to-speech (TTS) has made significant progress in replicating unseen voices, yet balancing generation quality and inference efficiency remains challenging. Autoregressive models suffer from high latency, whil…

  2. arXiv cs.AI TIER_1 English(EN) · Alef Iury Siqueira Ferreira, Lucas Rafael Stefanel Gris, Luiz Fernando de Ara\'ujo Vidal, Frederico Santos de Oliveira, Christopher Dane Shulby, Anderson da Silva Soares, Arlindo Rodrigues Galv\~ao Filho ·

    Mask, Sample, Revise: A Revisable CTMC Inference Stack for Guided Discrete Flow Matching Text-to-Speech

    arXiv:2606.13989v1 Announce Type: cross Abstract: Recent alignment-free non-autoregressive (NAR) text-to-speech (TTS) models formulate synthesis as a conditional infilling task, bypassing explicit duration predictors and external aligners. When speech is represented with neural c…