The Decoder·14d agoAI models' written reasoning steps correspond to distinct internal patterns, a new study finds#activation-analysis#chain-of-thought#interpretability 4 min2
arXiv cs.AI / cs.LG / cs.CL·19d agoYou Can't Prefer Emotions You Don't Sample: Intensity Undershoot in DPO-Tuned LLMs#dpo#emotion-control#llama-3.1-8bAI research
Hugging Face daily papers·24d agoSafety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal#llm#over-refusal#qwen3-8b1