The OpenMOSS-Team has released MOSS-TTS-Local-Transformer-v1.5, an updated version of their text-to-speech model. This new version builds upon its predecessor, v1.0, by offering higher-fidelity stereo audio modeling with a 48 kHz sampling rate. It also features improved multilingual synthesis with language tags, more stable voice cloning, better handling of long-reference audio for short text cloning, and more precise punctuation-following prosody. The model now supports explicit pause control using inline markers and expands its language support to 31 languages, including new additions like Cantonese, Dutch, and Hindi. AI
IMPACT Enhances capabilities for developers working with open-source TTS models, offering higher fidelity and broader language support.
RANK_REASON Release of a new version of an open-source text-to-speech model with feature improvements. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Trending Models →
- Google Colab
- Kaggle
- MOSS-Audio-Tokenizer-v2
- MOSS-TTS-Local-Transformer-v1.0
- OpenMOSS-Team
- OpenMOSS-Team/MOSS-TTS-Local-Transformer-v1.5
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →