PulseAugur
EN
LIVE 09:32:56

Marco-Voice system integrates voice cloning and emotion control for expressive speech synthesis

Researchers have developed Marco-Voice, a novel speech synthesis system that integrates voice cloning with emotional control. This system addresses challenges in generating expressive, natural speech while maintaining speaker identity across different emotions and languages. Marco-Voice utilizes a speaker-emotion disentanglement mechanism and a rotational emotional embedding integration method for fine-grained control. The system was evaluated using the CSEMOTIONS dataset, a newly constructed Mandarin speech dataset, and demonstrated competitive performance in objective and subjective metrics for speech clarity and emotional richness. AI

IMPACT This research advances expressive neural speech synthesis, potentially enabling more natural and emotionally nuanced AI-driven voice applications.

RANK_REASON The cluster contains a technical report detailing a new speech synthesis system, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Marco-Voice system integrates voice cloning and emotion control for expressive speech synthesis

COVERAGE [1]

  1. arXiv cs.CL TIER_1 Română(RO) · Fengping Tian, Chenyang Lyu, Xuanfan Ni, Haoqin Sun, Qingjuan Li, Zhiqiang Qian, Haijun Li, Longyue Wang, Zhao Xu, Weihua Luo, Kaifu Zhang ·

    Marco-Voice Technical Report

    arXiv:2508.02038v5 Announce Type: replace Abstract: This paper presents a multifunctional speech synthesis system that integrates voice cloning and emotion control speech synthesis within a unified framework. The goal of this work is to address longstanding challenges in achievin…