arXiv cs.AI / cs.LG / cs.CL·2d agoTo Trust or Not to Trust: Retrieval-Augmented Fact Checking in Speech#verispeak#fact-checking#speechAI research
arXiv cs.AI / cs.LG / cs.CL·2d agoDo Audio Language Models Hear and Read Distinctive Features Alike?#audio-language-models#qwen2.5-omni#phonologyAI research
MarkTechPost·3d agoNVIDIA Releases Nemotron 3 Diarization: A 100M-Parameter Open-Weight Model That Tracks 8 Speakers in Real Time#nvidia#nemotron#diarization 2 sources 5 min
The Decoder·3d agoAlibaba launches Qwen Audio 3.1 with new models and slashes AI audio prices by up to 95 percent#alibaba#qwen#qwen-audio 9 sources
MarkTechPost·4d agoKyutai Releases Voice of Reason: A Speech-Native Model that Solves Spoken Math with Reinforcement Learning#kyutai#voice-of-reason#glm-4-voice 4 min
arXiv cs.AI / cs.LG / cs.CL·5d agoToneCL: Contrastive Learning for Few-Shot Syllable-Level Tone Classification#contrastive-learning#few-shot#speechAI research
arXiv cs.AI / cs.LG / cs.CL·11d agoECHO: A Matched-Contrast Benchmark for Context-Sensitive Turn-Taking in Full-Duplex Dialogue#benchmark#chinese#dialogueAI research
arXiv cs.AI / cs.LG / cs.CL·15d agoContinue, Adapt, or Yield: In-Turn Adaptation to Overlapping Speech in Full-Duplex Agents#conversational-ai#evaluation#full-duplexAI research
arXiv cs.AI / cs.LG / cs.CL·15d agoMP-Bench: Evaluating Voice Agents as a Multiparty Conversation Participant#benchmark#conversational-ai#evaluationAI research2
The Decoder·16d agoOpenAI's GPT-Live-1 API lets developers build apps that talk and listen at the same time#full-duplex#gpt-live-1#openai1
OpenAI News·17d agoBuild more natural voice experiences with GPT‑Live‑1 in the API#api#full-duplex#gpt-live-1 5 min1
Hugging Face Blog·Aug 10, 2026Build Low-Latency Multilingual Voice Agents: Open Weights & Full Deployment Control with NVIDIA Magpie TTS#magpie-tts#nvidia#open-weightsModel release