Hugging Face daily papers·10d agoGeometry of Values: Task Vector Composition for Ethical Preference Alignment in Language Models#alignment#benchmark#dpo
OpenAI News·12d agoHow Fyxer built an AI executive assistant people trust#ai-assistant#dpo#fine-tuning 5 min
Hugging Face daily papers·16d agoStepAudio 3 Music Technical Report#diffusion-transformer#dpo#flow-matching
arXiv cs.AI / cs.LG / cs.CL·16d agoDomain-Specific Hallucination Detection in Large Language Models#calibration#deberta-v3#dpoAI research1
arXiv cs.AI / cs.LG / cs.CL·16d agoRetroThinker: Enabling Retrospective Thinking in Speech LLMs#chain-of-thought#dpo#gsm8kAI research1
arXiv cs.AI / cs.LG / cs.CL·19d agoYou Can't Prefer Emotions You Don't Sample: Intensity Undershoot in DPO-Tuned LLMs#dpo#emotion-control#llama-3.1-8bAI research
Hugging Face daily papers·27d agoPLC-DPO: Posterior Label Correction in Noisy and Ambiguous Preference Optimization#alignment#dpo#noisy-labels
Hugging Face daily papers·27d agoAgenticGen: Reward-Guided Agentic Video Generation for Advertising#advertising#agents#dpo