Leaky Students: Membership Inference against On-Policy Distillation
Researchers show on-policy distillation students leak teacher-record membership, with Leaky reaching mean AUROC 0.875.
On-policy distillation trains a student to match a teacher's next-token distributions, and data supplied to the teacher may be sensitive. Leaky samples fresh student trajectories and compares token log-probabilities with reference models trained without candidate records, applying a Leaky ReLU to the gaps. Across fifteen math, medical question-answering, and code targets, it reaches mean AUROC 0.875 versus 0.614 for the strongest baseline, or 0.826 for that baseline on the same trajectories. The results show OPD students can reveal membership of records used for teacher supervision.
- First systematic membership-inference study of on-policy distillation.
- Leaky compares fresh student trajectories with reference models lacking candidate records.
- Mean AUROC is 0.875 across fifteen math, medical, and code targets.
- Strongest baseline scores 0.614, or 0.826 on the same trajectories.
- Fixed reference-answer losses often miss sparse membership signals.
Full article222 words · extracted from arxiv.org · click to collapse
On-policy distillation (OPD) trains a student to match a teacher's next-token distributions on student-generated trajectories. However, privileged information supplied to the teacher for OPD training may contain sensitive data. Whether the student leaks private information about the records supplied to the teacher during distillation remains poorly understood. To the best of our knowledge, we present the first systematic study of membership inference in this setting. We find that fresh student trajectories expose sparse membership signals that fixed reference-answer losses often miss. These signals are mixed with probability changes caused by training on other records. We introduce Leaky, which samples fresh trajectories from the target model and compares its token log-probabilities with the maximum across matched reference models trained without the candidate records. It applies Leaky ReLU to the resulting gaps, preserving positive gaps and downweighting negative gaps as an approximate correction for incidental positive gaps in non-members. Across fifteen targets spanning mathematics, medical question answering, and code generation, Leaky outperforms all evaluated baselines and achieves mean AUROC 0.875, compared with 0.614 for the strongest baseline on each target in the main evaluation. On the same sampled trajectories, the strongest baseline achieves mean AUROC 0.826. These results show that students trained through OPD can expose the membership of records used for teacher supervision, even when fixed reference-answer losses provide little evidence of membership.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.33136