What Does Privileged Information Add to On-Policy Self-Distillation?
Study of on-policy self-distillation on 5,319 AMPLE-Math problems finds reference-free distillation drives most gains, with modest extra benefit from privileged references.
The authors build AMPLE-Math, a suite of 5,319 mathematical problems with six reasoning views sharing the same answers, to isolate the value of privileged references in on-policy self-distillation. With a thinking-enabled teacher supervising direct-response rollouts, reference-free distillation accounts for most of Qwen3-1.7B's improvement in domain and on external benchmarks; extra reference benefit is modest, strongest for polished solutions, while complete traces add two percentage points for SmolLM3-3B at step 50. Gains reverse when short direct-response rollouts are replaced with long thinking-enabled rollouts. The findings suggest OPSD improves access to existing reasoning capabilities via cross-mode parameter sharing, with references valued for what they add to that transfer.