PulseAugur
EN
LIVE 00:05:09

Alibaba Qwen launches full-duplex voice model Qwen-Audio-3.1-Realtime

Alibaba's Qwen team has launched Qwen-Audio-3.1, a suite of five audio models including a full-duplex speech model named Qwen-Audio-3.1-Realtime designed for voice agents that interact with tools. This new model boasts a 262K token context window and features like function calling and web search. The system operates on a "Think, Act, Speak" architecture, utilizing multiple models for decision-making, speech-to-text conversion, and text-to-speech rendering. Alibaba has also significantly reduced pricing for its audio models, with Qwen-Audio-3.1-Realtime seeing an approximately 85% price cut. AI

IMPACT Enhances voice agent capabilities with full-duplex interaction and tool-calling, potentially improving user experience in conversational AI applications.

RANK_REASON Frontier-lab model release with system card [lever_c_demoted from frontier_release: ic=1 ai=1.0]

Read on MarkTechPost →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Alibaba Qwen launches full-duplex voice model Qwen-Audio-3.1-Realtime

COVERAGE [1]

  1. MarkTechPost TIER_1 English(EN) · Asif Razzaq ·

    Alibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex Voice Model Trained to Think, Act, and Decide When to Speak

    <p>Alibaba's Qwen team released Qwen-Audio-3.1-Realtime, a full-duplex voice model trained to reason, call tools and decide when to speak. On a τ-Voice adaptation, task success rises to 82.0% from 78.4%. Replies to background speech drop from 73% to 13%. It is available now as an…