PulseAugur
EN
LIVE 09:33:11

Confucius4-TTS: New TTS System Enables Transcript-Free Cross-Lingual Voice Cloning

Researchers have developed Confucius4-TTS, a novel text-to-speech system capable of transcript-free, cross-lingual zero-shot voice cloning. This system supports 14 languages and utilizes a two-stage architecture comprising text-to-semantic and semantic-to-acoustic modules. The text-to-semantic module employs a learnable speaker encoder to capture timbre from speech representations, while the semantic-to-acoustic module generates mel-spectrograms. Confucius4-TTS demonstrates strong performance on cross-lingual benchmarks, achieving a 3.73% average Word Error Rate on CV3-Eval and outperforming other open-source and commercial systems in human evaluations. AI

IMPACT This advancement could significantly improve accessibility and efficiency in cross-lingual voice applications and content creation.

RANK_REASON Academic paper detailing a new TTS model with novel capabilities. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.CL →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Confucius4-TTS: New TTS System Enables Transcript-Free Cross-Lingual Voice Cloning

COVERAGE [1]

  1. arXiv cs.CL TIER_1 English(EN) · Huaxuan Wang, Huimin Wang, Ruiyu Zhang, Yingjie Li, Yitao Duan ·

    Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder

    arXiv:2608.11650v1 Announce Type: cross Abstract: Recent advances in zero-shot text-to-speech (TTS) have substantially improved speech quality and voice cloning fidelity. However, many zero-shot TTS systems still depend on audio prompt transcripts at inference time. This dependen…