Alibaba launches Qwen Audio 3.1 with new models and slashes AI audio prices by up to 95 percent
Alibaba's Qwen team released five Qwen-Audio-3.1 speech models and cut audio prices by up to 95 percent.
Alibaba's Qwen team launched Qwen-Audio-3.1, a five-model lineup for speech recognition, text-to-speech, and real-time conversation. The ASR model targets multilingual and dialect recognition and removes filler words, while ASR-Next adds speaker timestamps plus emotion, ambient-sound, and machine-noise detection. TTS supports prompted emotion, speed, and style, and TTS-Next combines a language model with diffusion to produce voice, effects, and background audio. On Qwen Cloud, TTS prices drop about 70 percent, Realtime about 85 percent, and ASR up to 95 percent.
48