Alibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex Voice Model Trained to Think, Act, and Decide When to Speak
Alibaba released API-only Qwen-Audio-3.1-Realtime, a full-duplex voice model for tool-using agents.
Alibaba's Qwen team released Qwen-Audio-3.1, a five-model audio stack for ASR, TTS, and realtime interaction, led by Qwen-Audio-3.1-Realtime, a full-duplex model for voice agents that call tools. The plus variant is live on QwenCloud over WebSocket with a 262K context, function calling, and web search; no open weights were announced. Listed pricing is $6.4 per million audio input tokens and $24 per million text-and-audio output tokens, alongside cuts of about 85% on Realtime, 70% on TTS, and up to 95% on ASR. Against version 3.0, Audio MultiChallenge rose from 47.12 to 52.21, the 14-language BBA average from 81.7% to 88.1%, and FLEURS WER fell from 9.01 to 3.98.