Researchers have developed Marco-Voice, a novel speech synthesis system that integrates voice cloning with emotional control. This system addresses challenges in generating expressive, natural speech while maintaining speaker identity across different emotions and languages. Marco-Voice utilizes a speaker-emotion disentanglement mechanism and a rotational emotional embedding integration method for fine-grained control. The system was evaluated using the CSEMOTIONS dataset, a newly constructed Mandarin speech dataset, and demonstrated competitive performance in objective and subjective metrics for speech clarity and emotional richness. AI
IMPACT This research advances expressive neural speech synthesis, potentially enabling more natural and emotionally nuanced AI-driven voice applications.
RANK_REASON The cluster contains a technical report detailing a new speech synthesis system, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- Chenyang Lyu
- CSEMOTIONS
- DagsHub
- Gotit.pub
- Hugging Face
- Marco-Voice
- ScienceCast
- Standard Chinese
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →