Alibaba's Qwen team has launched Qwen-Audio-3.1, a suite of five audio models including a full-duplex speech model named Qwen-Audio-3.1-Realtime designed for voice agents that interact with tools. This new model boasts a 262K token context window and features like function calling and web search. The system operates on a "Think, Act, Speak" architecture, utilizing multiple models for decision-making, speech-to-text conversion, and text-to-speech rendering. Alibaba has also significantly reduced pricing for its audio models, with Qwen-Audio-3.1-Realtime seeing an approximately 85% price cut. AI
IMPACT Enhances voice agent capabilities with full-duplex interaction and tool-calling, potentially improving user experience in conversational AI applications.
RANK_REASON Frontier-lab model release with system card [lever_c_demoted from frontier_release: ic=1 ai=1.0]
- Alibaba Group
- Core-Cocktail SFT
- Full-Duplex-Bench v1.5
- Full-Duplex-Bench v3.0
- Grpo
- M²-OPD
- Qwen
- Qwen-Audio-3.1
- Qwen-Audio-3.1-ASR-Flash-Filetrans
- Qwen-Audio-3.1-Realtime
- QwenCloud
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →