Two new research papers explore advancements in zero-shot text-to-speech (TTS) technology, focusing on discrete flow matching techniques. The first paper introduces DiFlow-TTS, a framework that uses a discrete flow matching approach to balance generation quality and inference efficiency, addressing limitations of autoregressive and continuous-space flow-based models. The second paper, "Mask, Sample, Revise," proposes an inference-time stack for discrete flow matching TTS, enhancing control and robustness in generating speech from neural codec tokens without explicit duration predictors. AI
IMPACT These papers introduce novel techniques for text-to-speech synthesis, potentially leading to more efficient and higher-quality voice generation systems.
RANK_REASON Two academic papers published on arXiv detailing new methods for text-to-speech synthesis.
- Alef Iury Siqueira Ferreira
- Continuous-Time Markov Chain
- Discrete Flow Matching
- Mask, Sample, Revise
- Text-to-Speech
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- DiFlow-TTS
- Gotit.pub
- Hugging Face
- Ngoc Son Nguyen
- ScienceCast
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →