arXiv cs.AI / cs.LG / cs.CL·3d agoWhen and Where to Trust the Teacher: Unifying On-Policy Distillation and GRPO through Entropy-Calibrated Credit Assignment#grpo#on-policy-distillation#reinforcement-learningAI research
Hugging Face daily papers·6d ago1% of Tokens Can Be Enough: On Gradient Estimation in On-Policy Distillation#on-policy-distillation#gradient-estimation#token-selection
Hugging Face daily papers·9d agoCalibrating Teacher--Student Discrepancy for On-Policy Distillation#distillation#knowledge-distillation#on-policy-distillation
Hugging Face daily papers·10d agoWhen EOS Tokens Disagree: Understanding Length Inflation in On-Policy Distillation#distillation#eos-tokens#gemma
Hugging Face daily papers·10d agoWhat Does Privileged Information Add to On-Policy Self-Distillation?#math-benchmarks#on-policy-distillation#qwen31
Hugging Face daily papers·10d agoRetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning#agentic-rl#alfworld#distillation
Hugging Face daily papers·14d agoLightning Weave: Improving the Accuracy-Efficiency Frontier of Reasoning Models through Capability Composition#capability-composition#dopd#efficient-reasoning
Hugging Face daily papers·19d agoNeoHorse-1: Towards Recursive Self-Improvement via Agentic Post-Training with Routing Harness#agentic-post-training#benchmarks#fine-tuning1
Hugging Face daily papers·25d agoVerify Before You Distill: Prompt-Level Teacher Gating for On-Policy Distillation#grpo#llm#on-policy-distillation