Sharpen Without Search: On-Policy Distillation of Sequence-Level Power Distribution
On-policy power distillation (OPPD) trains a language model to emit sharpened answers in a single generation, lifting MATH500 by up to 23 points and beating same-budget GRPO without multi-candidate search.
On-policy power distillation (OPPD) trains a language model to emit answers from a sharpened sequence-level power distribution in one generation, replacing multi-candidate power sampling. A sequential Monte Carlo sampler has the student propose candidates while a frozen teacher's power distribution weights a maximum-likelihood update. Compared with the untrained model at the same temperature, single-generation accuracy rises by up to 23.0 points on MATH500 and 27.3 points on GSM8K, and a single sample exceeds published 64-candidate power sampling by 2.4 and 3.5 points. Against GRPO trained from the same checkpoint with the same budget, OPPD scores 3.8, 4.0 and 5.4 points higher on MATH500, GSM8K and AIME without requiring reference answers, and applying OPPD after GRPO adds up to 9.3 points. A single loss coefficient sets the absorbed sharpening exponent between 1.19 and 2.02, and math-only OPPD training also raises HumanEval accuracy by up to 5.3 points.
- OPPD trains a student model to sample from a sharpened sequence-level power distribution in one generation, replacing multi-candidate power search.
- Accuracy rises by up to 23.0 points on MATH500 and 27.3 points on GSM8K versus the untrained model at the same temperature.
- A single OPPD sample exceeds published 64-candidate power sampling by 2.4 points on MATH500 and 3.5 points on GSM8K.
- OPPD beats GRPO trained from the same checkpoint with the same budget by 3.8 (MATH500), 4.0 (GSM8K) and 5.4 (AIME) points, without reference answers.
- Applying OPPD after GRPO adds up to 9.3 points.
- A single loss coefficient controls the absorbed sharpening exponent, which ranges from 1.19 to 2.02.
- Math-only OPPD training lifts HumanEval accuracy by up to 5.3 points.
Coverage timelineoldest first · each row is one article
- · 4d agoSharpen Without Search: On-Policy Distillation of Sequence-Level Power Distribution
Hugging Face daily papers· 58
On-policy power distillation trains one-shot sharpened reasoning, lifting MATH500 by up to 23 points.
- · 3d agoSharpen Without Search: On-Policy Distillation of Sequence-Level Power Distribution
arXiv cs.AI / cs.LG / cs.CL· 55
On-policy power distillation trains models to sample sharpened answers once, lifting MATH500 by up to 23 points.