Alibaba launches Qwen Audio 3.1 with new models and slashes AI audio prices by up to 95 percent
Alibaba's Qwen team released five Qwen-Audio-3.1 speech models and cut audio prices by up to 95 percent.
Alibaba's Qwen team launched Qwen-Audio-3.1, a five-model lineup for speech recognition, text-to-speech, and real-time conversation. The ASR model targets multilingual and dialect recognition and removes filler words, while ASR-Next adds speaker timestamps plus emotion, ambient-sound, and machine-noise detection. TTS supports prompted emotion, speed, and style, and TTS-Next combines a language model with diffusion to produce voice, effects, and background audio. On Qwen Cloud, TTS prices drop about 70 percent, Realtime about 85 percent, and ASR up to 95 percent.
- Five Qwen-Audio-3.1 models cover ASR, TTS, and realtime speech.
- ASR-Next timestamps speakers and detects emotion and ambient noise.
- TTS-Next generates voice, effects, and background audio together.
- ASR pricing falls by up to 95 percent on Qwen Cloud.
Full article200 words · extracted from the-decoder.com · click to collapse
Alibaba's AI team Qwen has released Qwen-Audio-3.1, a lineup of five models for speech recognition (ASR), text-to-speech (TTS), and real-time interaction. The ASR model improves multilingual and dialect recognition and automatically cleans up filler words and repetitions. ASR-Next adds multi-speaker identification with timestamps and detects emotions, ambient sounds, and machine noise. TTS handles multilingual synthesis with natural cross-language voice transfer. Users control emotion, speed, and style through simple text prompts like "Read this with a sharp, commanding tone, demanding respect."
TTS-Next pairs a language model with a diffusion approach to generate voice, sound effects, and background audio in a single pass. The real-time model supports simultaneous speaking and listening with instant interruption. When it detects a low mood, it responds more slowly and with more empathy, according to Qwen. Alibaba is also slashing prices. TTS drops about 70 percent, Realtime roughly 85 percent, and ASR up to 95 percent. More details on the blog and on Qwen Cloud.
AI News Without the Hype – Curated by Humans
Subscribe to THE DECODER for ad-free reading, a weekly AI newsletter, our exclusive "AI Radar" frontier report six times a year, full archive access, and access to our comment section.
Text extracted automatically; images, tables and formatting may be missing. Original: https://the-decoder.com/alibaba-launches-qwen-audio-3-1-with-five-new-models-and-slashes-ai-audio-prices-by-up-to-95-percent/