Researchers have developed a method for creating high-quality Text-to-Speech (TTS) for low-resource languages, specifically focusing on Modern Greek. Their approach involves curating audiobook data using WhisperX for alignment and filtering, then fine-tuning the Parler-TTS model. To address speaker drift issues caused by LLM-generated prompts, they implemented deterministic prompts and a LoRA stage for speaker identity, achieving a Word Error Rate of 10.7% and a Mean Opinion Score for Intelligibility of 4.00. AI
IMPACT Demonstrates a viable approach for developing high-quality TTS in languages with scarce data, potentially improving accessibility.
RANK_REASON Academic paper detailing a new methodology for TTS in a low-resource language. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →