An RL View of OPD: Least Square Policy Distillation for Sample-Efficient LLM Reasoning
Least-Square Policy Distillation beats on-policy distillation baselines by 1.59 Avg@16 on math reasoning.
The paper links on-policy distillation's reverse-KL objective to KL-regularized policy optimization and introduces Least-Square Policy Distillation. LSPD adds optimistic exploration and off-policy reuse of earlier trajectories while aiming to keep policy diversity. Theory ties the method to optimistic value-based learning and gives an idealized O(log K) regret bound under online exploration. Across six math reasoning benchmarks it gains +1.59 Avg@16 on average, stays stronger as Pass@k grows to 64, and an off-policy variant matches vanilla OPD using only the first 25% of rollout batches.
- LSPD connects reverse-KL on-policy distillation to KL-regularized RL.
- Idealized analysis gives an O(log K) online regret bound.
- Average gain is +1.59 Avg@16 across six math benchmarks.
- Off-policy variant matches OPD with the first 25% of rollouts.
- Pass@k through k=64 shows better preserved policy diversity.
Full article185 words · extracted from huggingface.co · click to collapse
We study on-policy distillation (OPD) through the lens of reinforcement learning, establishing a connection between the reverse-KL objective in OPD and KL-regularized policy optimization. Building on this connection, we introduce Least-Square Policy Distillation (LSPD), an RL-inspired framework that brings optimistic exploration and off-policy data reuse from value-based RL into policy distillation. LSPD preserves policy diversity through exploration while improving rollout efficiency by repeatedly learning from previously collected trajectories. Our theoretical analysis connects LSPD to optimistic value-based learning and shows that its idealized formulation achieves a sharp mathcal O(log K) regret bound under online exploration. Empirically, LSPD consistently outperforms existing distillation baselines across six mathematical reasoning benchmarks and diverse teacher-student settings, with average gains of +1.59 points in Avg@16. Remarkably, through Pass@k evaluations up to k=64, we found that LSPD better preserves policy diversity by achieving stronger performance as k grows. Its fully off-policy variant achieves comparable performance to vanilla OPD using only the first 25% of rollout batches. Together, these results provide an RL perspective on OPD that offers both a principled interpretation and a practical route toward more effective and rollout-efficient language model distillation.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.35505