PulseAugur
EN
LIVE 02:29:16
中文(ZH) 阿里巴巴发布语音合成大模型Qwen-Audio-3.0-TTS

Alibaba launches Qwen-Audio-3.0-TTS with 16-language support and top leaderboard ranking

Alibaba's Tongyi Lab has launched Qwen-Audio-3.0-TTS, an advanced text-to-speech model available in two tiers: Flash for real-time interaction and Plus for high-quality generation. The model boasts significant improvements in fine-grained control, instruction following, multilingual support across 16 languages, and robustness. Qwen-Audio-3.0-TTS-Plus has achieved the top rank on the Artificial Analysis leaderboard for text-to-speech quality, with its preview version also previously holding the top spot. AI

IMPACT Sets new SOTA on TTS quality benchmarks and offers advanced control features for developers.

RANK_REASON Frontier-lab model release with system card [lever_c_demoted from frontier_release: ic=2 ai=1.0]

Read on 36氪 (36Kr) →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

Alibaba launches Qwen-Audio-3.0-TTS with 16-language support and top leaderboard ranking

COVERAGE [2]

  1. 36氪 (36Kr) TIER_1 中文(ZH) ·

    Alibaba releases speech synthesis large model Qwen-Audio-3.0-TTS

    36氪获悉,7月20日,阿里巴巴发布语音合成大模型Qwen-Audio-3.0-TTS,在细粒度标签控制、freestyle指令遵循、多语种与方言覆盖以及复杂声学鲁棒性等方面系统性提升,包含面向实时交互的Flash版本(首包延时300ms级别)和面向高质量生成的Plus版本,让合成语音从“能说话”走向“会表达”。目前,在全球第三方榜单Artificial Analysis,Qwen-Audio-3.0-TTS-Plus登顶,其预览版Fun-Realtime-TTS也曾在一个月前登顶该榜单。

  2. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    Alibaba’s Tongyi Lab Releases Qwen-Audio-3.0-TTS, a Hosted Text-to-Speech Model in Flash and Plus Tiers Across 16 Languages

    <p>Alibaba&#8217;s Tongyi Lab has released Qwen-Audio-3.0-TTS, a production-oriented text-to-speech (TTS) system. The model ships in two variants from the same lineage. Flash targets real-time interaction. Plus targets high-quality generation. Both are delivered as hosted models …