Hugging Face daily papers·9d agoOmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue#audio-visual-dialogue#benchmark#data-synthesis
arXiv cs.AI / cs.LG / cs.CL·9d agoMulti-Dimensional Prosody Judgment For Live Streaming Speech Synthesis#distillation#grpo#llm-judgeAI research1