PulseAugur
EN
LIVE 13:53:08
中文(ZH) 阿里巴巴发布语音合成大模型Qwen-Audio-3.0-TTS

Alibaba's Qwen-Audio-3.0-TTS leads text-to-speech rankings

Alibaba has launched Qwen-Audio-3.0-TTS, a new text-to-speech model that offers significant improvements in control, multilingual support, and acoustic robustness. The model comes in two versions: Flash for real-time interaction with low latency and Plus for high-quality generation. Qwen-Audio-3.0-TTS-Plus has achieved the top position on the Artificial Analysis Speech Arena leaderboard, outperforming competitors like Google and Speechify. AI

IMPACT Sets a new benchmark for TTS quality and control, potentially influencing enterprise adoption and further research in expressive speech synthesis.

RANK_REASON New model release from a major tech company (Alibaba) with performance claims and benchmark results.

Read on 36氪 (36Kr) →

AI-generated summary · Google Gemini · from 4 sources. How we write summaries →

Alibaba's Qwen-Audio-3.0-TTS leads text-to-speech rankings

COVERAGE [4]

  1. 36氪 (36Kr) TIER_1 中文(ZH) ·

    Alibaba releases speech synthesis large model Qwen-Audio-3.0-TTS

    36氪获悉,7月20日,阿里巴巴发布语音合成大模型Qwen-Audio-3.0-TTS,在细粒度标签控制、freestyle指令遵循、多语种与方言覆盖以及复杂声学鲁棒性等方面系统性提升,包含面向实时交互的Flash版本(首包延时300ms级别)和面向高质量生成的Plus版本,让合成语音从“能说话”走向“会表达”。目前,在全球第三方榜单Artificial Analysis,Qwen-Audio-3.0-TTS-Plus登顶,其预览版Fun-Realtime-TTS也曾在一个月前登顶该榜单。

  2. The Decoder TIER_1 English(EN) · Matthias Bastian ·

    Alibaba's Qwen Audio 3.0 TTS Plus tops the competition in the text-to-speech rankings

    <p><img alt="" class="attachment-full size-full wp-post-image" height="1152" src="https://the-decoder.com/wp-content/uploads/2026/07/qwen_logo-1.png" style="height: auto; margin-bottom: 10px;" width="2048" /></p> <p> Alibaba's Qwen Audio 3.0 TTS Plus tops Artificial Analysis' Spe…

  3. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    Alibaba’s Tongyi Lab Releases Qwen-Audio-3.0-TTS, a Hosted Text-to-Speech Model in Flash and Plus Tiers Across 16 Languages

    <p>Alibaba&#8217;s Tongyi Lab has released Qwen-Audio-3.0-TTS, a production-oriented text-to-speech (TTS) system. The model ships in two variants from the same lineage. Flash targets real-time interaction. Plus targets high-quality generation. Both are delivered as hosted models …

  4. Mastodon — mastodon.social TIER_1 Polski(PL) · aisight ·

    Alibaba Cloud's Qwen-Audio-3.0-TTS-Plus model dominated the Speech Arena ranking, beating solutions from Google and Speechify. The new leader offers unprecedented control

    Model Qwen-Audio-3.0-TTS-Plus od Alibaba Cloud zdominował ranking Speech Arena, pokonując rozwiązania Google i Speechify. Nowy lider oferuje niespotykaną kontrolę nad emocjami, choć jego wysoka jakość wiąże się z ceną segmentu premium. # si # ai # sztucznainteligencja # wiadomośc…