arXiv cs.AI / cs.LG / cs.CL·3d agoOrder-Invariant Answers, Order-Sensitive Representations in Mathematical Reasoning#language-models#mathematical-reasoning#representationsAI research
arXiv cs.AI / cs.LG / cs.CL·3d agoWhen and Where to Trust the Teacher: Unifying On-Policy Distillation and GRPO through Entropy-Calibrated Credit Assignment#grpo#on-policy-distillation#reinforcement-learningAI research
Hugging Face daily papers·18d agoAn Open Recipe for IMO Gold: Training Nemotron for Olympiad Mathematics#mathematical-reasoning#nemotron#olympiad-mathematics