Researchers have developed a new method called Speech-Rewarded Style Planning (SRSP) to improve controllable text-to-speech (TTS) systems. Unlike previous approaches that rely on text descriptions of speech style, SRSP trains a style planner using a frozen TTS model. This planner generates instructions that are optimized using group-relative policy optimization (GRPO) with the likelihood of target speech tokens as the reward. Experiments on the ISCSLP 2026 CoT-TTS corpus showed that SRSP outperforms baseline methods in terms of speech-style similarity, emotion similarity, and mel-cepstral distortion, while also demonstrating improved contextual appropriateness and reference consistency in expressive speech evaluations. AI
IMPACT Improves control and naturalness in text-to-speech systems, potentially leading to more expressive and contextually appropriate synthesized speech.
RANK_REASON The cluster contains a research paper detailing a new method for text-to-speech synthesis. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- CatalyzeX Code Finder for Papers
- DagsHub
- Gotit.pub
- Grpo
- Hugging Face
- Influence Flower
- ISCSLP 2026 CoT-TTS corpus
- ScienceCast
- Speech-Rewarded Style Planning (SRSP)
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →