arXiv cs.AI / cs.LG / cs.CL·4d agoTrain Where the Quantized Model Goes: On-Policy Distillation for Low-Bit Reasoning#quantization#distillation#on-policyAI research
Hugging Face daily papers·5d agoonPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction#onpanda#alignment#annotation 2 sources
Hugging Face daily papers·16d agoMInTRL: Off-policy Intervention can boost On-policy RL#advantage-regression#exploration#on-policy