PulseAugur
EN
LIVE 06:51:11

New TTS Model Learns Expressive Speech from Open-Ended Instructions

Researchers have developed Poly-InstructTTS, a novel text-to-speech model capable of generating expressive speech based on natural language instructions. The system utilizes a large, in-the-wild audiovisual dataset of 1,000 hours, annotated with over 1,000 distinct emotions and styles. Poly-InstructTTS employs a prompt-free GPT architecture and a flow-matching module for timbre injection, along with a speaker fine-tuning procedure to maintain persona. Evaluations indicate strong performance in instruction adherence and expressiveness, with accompanying audio demos and an expanded test set available. AI

IMPACT Enables more nuanced and controllable expressive speech generation, potentially impacting voice assistants and content creation.

RANK_REASON Academic paper detailing a new model and dataset for speech synthesis. [lever_c_demoted from research: ic=1 ai=1.0]

Read on arXiv cs.AI →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

New TTS Model Learns Expressive Speech from Open-Ended Instructions

COVERAGE [1]

  1. arXiv cs.AI TIER_1 English(EN) · Junhui Zhang, Qianhui Xu, Qingxiang Guo, Dawei Yang, Ling Miao, Qiangqiang Wang, Yang Song ·

    Poly-InstructTTS: Learning In-the-Wild Expressive Speech Synthesis from Open-Ended Instructions

    arXiv:2608.20387v1 Announce Type: cross Abstract: While recent text-to-speech (TTS) models achieve high naturalness, controlling fine-grained expression via natural-language instructions remains challenging. We introduce Poly- InstructTTS, which learns expressive speech from open…