Alibaba's Tongyi Lab has launched Qwen-Audio-3.0-TTS, an advanced text-to-speech model available in two tiers: Flash for real-time interaction and Plus for high-quality generation. The model boasts significant improvements in fine-grained control, instruction following, multilingual support across 16 languages, and robustness. Qwen-Audio-3.0-TTS-Plus has achieved the top rank on the Artificial Analysis leaderboard for text-to-speech quality, with its preview version also previously holding the top spot. AI
IMPACT Sets new SOTA on TTS quality benchmarks and offers advanced control features for developers.
RANK_REASON Frontier-lab model release with system card [lever_c_demoted from frontier_release: ic=2 ai=1.0]
- Alibaba Group
- Artificial Analysis
- Fun-Realtime-TTS
- Qwen-Audio-3.0-TTS
- Qwen-Audio-3.0-TTS-Plus
- Alibaba Cloud Model Studio
- DashScope
- Tongyi Lab
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →