PulseAugur
EN
LIVE 23:39:49
中文(ZH) 长音频不丢词,行业词不用教,阿里发布Qwen-Audio-3.0-ASR-Flash

Alibaba launches Qwen-Audio-3.0-ASR-Flash with enhanced context and industry term recognition

Alibaba has launched Qwen-Audio-3.0-ASR-Flash, an upgraded speech recognition model. This new version enhances context consistency for long audio, improves industry-specific term recognition without manual lists, and offers more accurate customization for frequently used terms. Additionally, the model can now perform speech refinement, removing filler words and structuring spoken language into coherent text, all within a single step. AI

IMPACT This release enhances the accuracy and efficiency of speech-to-text applications, particularly for long-form content and specialized industries.

RANK_REASON Frontier-lab model release with system card. [lever_c_demoted from frontier_release: ic=1 ai=1.0]

Read on 雷峰网 (Leiphone) →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Alibaba launches Qwen-Audio-3.0-ASR-Flash with enhanced context and industry term recognition

COVERAGE [1]

  1. 雷峰网 (Leiphone) TIER_1 中文(ZH) ·

    Long audio without word loss, industry terms require no teaching, Alibaba releases Qwen-Audio-3.0-ASR-Flash

    <p>7月31日,阿里巴巴发布语音识别大模型Qwen-Audio-3.0-ASR-Flash,在上下文一致性、行业词识别和热词定制化三个维度系统性升级,同时具备语音润色能力,可直接输出结构化文本。</p><p>目前,Qwen-Audio-ASR系列已在会议纪要整理、实时字幕、教育录播、智能客服等场景中获得广泛验证,离线版本曾在全球权威AI评测平台Artificial Analysis以1.7%的错字率位列全球第一。识准专业词汇,是ASR模型走向行业落地的最后一公里,也是本次升级的核心。全新升级的3.0版本,聚焦通用的语音识别能力,不止于把每个字听对,更要…