Researchers have developed a novel approach to Text-to-Speech (TTS) synthesis that mimics how voice actors adjust their delivery based on performance directions. This method, termed direction-following TTS, aims to generate speech reflecting specific stylistic changes while maintaining the original speaker's identity and linguistic content. To overcome the scarcity of relevant training data, the team devised a scalable pipeline for constructing pseudo-triplets, which consist of a reference utterance, a direction text, and a modified utterance. This pipeline utilizes an impression-controllable TTS model for style variations and an LLM to generate natural language directions from estimated impression differences. Experiments show that these pseudo-triplets alone facilitate stable, speaker-preserving modifications, with further improvements in direction alignment and speaker similarity achieved when combined with recorded data. AI
IMPACT This research could lead to more nuanced and controllable speech synthesis, enabling applications that require dynamic vocal adjustments.
RANK_REASON Academic paper detailing a new method for TTS synthesis. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Scalable Direction-Following TTS via Voice Impression-Guided Pseudo Triplet Construction
- Text To Speech
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →