Alibaba Releases Qwen-Audio-3.1 and Cuts Audio Prices
Alibaba released five Qwen-Audio-3.1 speech models, including an API-only full-duplex realtime agent, and cut audio prices by up to 95%.
Alibaba's Qwen team released Qwen-Audio-3.1, a five-model audio stack for speech recognition, text-to-speech, and realtime conversation. The ASR model targets multilingual and dialect recognition and strips filler words, while ASR-Next adds speaker timestamps and detects emotion, ambient sound, and machine noise; TTS accepts prompted emotion, speed, and style, and TTS-Next pairs a language model with diffusion to generate voice, effects, and background audio together. The lineup is led by Qwen-Audio-3.1-Realtime, a full-duplex model for tool-using voice agents; MarkTechPost says the plus variant is live on Qwen Cloud over WebSocket with a 262K context, function calling, and web search, with no open weights announced, and lists $6.4 per million audio input tokens and $24 per million text-and-audio output tokens. Both reports say Qwen Cloud prices fell about 70 percent for TTS, about 85 percent for Realtime, and up to 95 percent for ASR. Against version 3.0, only MarkTechPost reports Audio MultiChallenge rising from 47.12 to 52.21, the 14-language BBA average from 81.7% to 88.1%, and FLEURS word error rate falling from 9.01 to 3.98, plus multi-turn attack success of 26.0% in Chinese and 23.5% in English. The sources do not disagree on the shared price cuts; the later report adds model, dollar-price, and benchmark detail the earlier one omits.
- Alibaba's Qwen team launched Qwen-Audio-3.1, a five-model lineup for ASR, TTS, and realtime conversation.
- ASR targets multilingual and dialect recognition and removes filler words; ASR-Next adds speaker timestamps and detects emotion, ambient sound, and machine noise.
- TTS supports prompted emotion, speed, and style; TTS-Next combines a language model with diffusion to produce voice, effects, and background audio.
- Qwen-Audio-3.1-Realtime is a full-duplex, API-only voice model; the plus variant is on Qwen Cloud over WebSocket with a 262K context, function calling, and web search, and no open weights were announced.
- Listed pricing is $6.4 per million audio input tokens and $24 per million text-and-audio output tokens; cuts are about 70% for TTS, about 85% for Realtime, and up to 95% for ASR.
- Versus version 3.0, Audio MultiChallenge rose from 47.12 to 52.21, the 14-language BBA average from 81.7% to 88.1%, and FLEURS WER fell from 9.01 to 3.98.
- Training uses Think, Act, and Speak layers, including GRPO in tool environments; reported multi-turn attack success is 26.0% in Chinese and 23.5% in English.
Coverage timelineoldest first · each row is one article
- · 6d agoAlibaba launches Qwen Audio 3.1 with new models and slashes AI audio prices by up to 95 percent
The Decoder· 48
Alibaba's Qwen team released five Qwen-Audio-3.1 speech models and cut audio prices by up to 95 percent.
- · 21h agoAlibaba Qwen Releases Qwen-Audio-3.1-Realtime: A Full-Duplex Voice Model Trained to Think, Act, and Decide When to Speak
MarkTechPost· 64
Alibaba released API-only Qwen-Audio-3.1-Realtime, a full-duplex voice model for tool-using agents.