Researchers have developed Poly-InstructTTS, a novel text-to-speech model capable of generating expressive speech based on natural language instructions. The system utilizes a large, in-the-wild audiovisual dataset of 1,000 hours, annotated with over 1,000 distinct emotions and styles. Poly-InstructTTS employs a prompt-free GPT architecture and a flow-matching module for timbre injection, along with a speaker fine-tuning procedure to maintain persona. Evaluations indicate strong performance in instruction adherence and expressiveness, with accompanying audio demos and an expanded test set available. AI
IMPACT Enables more nuanced and controllable expressive speech generation, potentially impacting voice assistants and content creation.
RANK_REASON Academic paper detailing a new model and dataset for speech synthesis. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- generative pre-trained transformer
- Gotit.pub
- Hugging Face
- InstructTTSEval
- Poly-InstructTTS
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →