Alibaba's Qwen team has introduced Qwen-Audio-3.0-ASR, a new Mixture-of-Experts large language model-based automatic speech recognition system. This model is designed to improve real-world utility by handling diverse dialects, dynamic entities, and long-range context. It supports transcription in 30 languages and 16 Chinese dialects, offering features like industry-domain entity recognition and hierarchical hotword customization. A streaming variant, Qwen-Audio-3.0-ASR-Streaming, is also available for low-latency applications, with evaluations showing competitive performance against systems like GPT-4o Transcribe and Gemini 3.1 Pro. AI
IMPACT Enhances real-world speech recognition capabilities, potentially improving applications requiring dialectal understanding and long-context processing.
RANK_REASON Publication of a technical report on a new ASR system. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →