Researchers have developed Confucius4-TTS, a novel text-to-speech system capable of transcript-free, cross-lingual zero-shot voice cloning. This system supports 14 languages and utilizes a two-stage architecture comprising text-to-semantic and semantic-to-acoustic modules. The text-to-semantic module employs a learnable speaker encoder to capture timbre from speech representations, while the semantic-to-acoustic module generates mel-spectrograms. Confucius4-TTS demonstrates strong performance on cross-lingual benchmarks, achieving a 3.73% average Word Error Rate on CV3-Eval and outperforming other open-source and commercial systems in human evaluations. AI
IMPACT This advancement could significantly improve accessibility and efficiency in cross-lingual voice applications and content creation.
RANK_REASON Academic paper detailing a new TTS model with novel capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →