Researchers have developed Confucius4-TTS, a new multilingual zero-shot text-to-speech system capable of cloning voices across 14 languages without needing transcripts of the reference audio. The system employs a two-stage architecture, with a text-to-semantic module utilizing a learnable speaker encoder and a semantic-to-acoustic module for mel-spectrogram generation. Confucius4-TTS demonstrates strong performance on cross-lingual benchmarks, achieving a 3.73% Word Error Rate on CV3-Eval and ranking highly in human evaluations. AI
IMPACT This development could significantly improve cross-lingual voice cloning capabilities for AI applications, enabling more natural and diverse voice synthesis.
RANK_REASON The cluster describes a technical report detailing a new TTS model release with associated code and checkpoints.
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →