The Decoder·3d agoAlibaba launches Qwen Audio 3.1 with new models and slashes AI audio prices by up to 95 percent#alibaba#qwen#qwen-audio 9 sources
MarkTechPost·6d agoBest Voice Cloning APIs in 2026: Speaker Similarity, Consent Checks, and Price per 1M Characters#cartesia#elevenlabs#fish-audio 6 min
arXiv cs.AI / cs.LG / cs.CL·8d agoNemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities#speech-to-speech#full-duplex#tool-callingAI research
arXiv cs.AI / cs.LG / cs.CL·9d agoMTVA-Bench: Evaluating the Language Model Inside Cascaded Voice Agents#asr#benchmark#llm-evaluationAI research
arXiv cs.AI / cs.LG / cs.CL·9d agoMulti-Dimensional Prosody Judgment For Live Streaming Speech Synthesis#distillation#grpo#llm-judgeAI research1
arXiv cs.AI / cs.LG / cs.CL·11d agoLACE: Layer-Wise Compression for Dynamic Frame Rate Codecs#audio-codec#compression#neural-codecAI research1
Hacker News · security·12d agoShow HN: Nari Qwen3-TTS and Qwen3-ASR – High accuracy, low latency and cost#asr#benchmarks#covalAI tools & infra 3 min1
Hugging Face daily papers·16d agoStepAudio 3 Gen Technical Report#audio-generation#autoregressive#rvq
Hugging Face daily papers·24d agoBuilding and Evaluating Fixed-Voice Thai TTS from Synthetic Speech#distillation#low-resource#speech-synthesis
Hugging Face Blog·Aug 10, 2026Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS#magpie-tts#nvidia#open-weightsModel release