Researchers have developed DriftTTS, a novel few-step text-to-speech model that achieves competitive synthesis quality without relying on distillation from pretrained teachers or adversarial discrimination. The model utilizes a distribution-matching drift objective in a mel-domain feature space, trained using on-policy rollout. On the LJSpeech dataset, DriftTTS demonstrated strong performance with low MCD and WER, and in blind listening tests, it achieved a MOS score comparable to ground truth and superior to Matcha-TTS. AI
IMPACT This research offers a new approach to few-step TTS synthesis, potentially reducing computational requirements and complexity in training.
RANK_REASON The cluster contains a research paper detailing a new model release. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →