PMOPD: Task Ordering, Cycling, and Parameter-Update Subspace Protection in Multi-Teacher On-Policy Distillation
PMOPD reduces cross-task interference in multi-teacher distillation, improving Code, Reason, and Math on Qwen and Llama.
Researchers propose PMOPD, which records low-dimensional subspace memories of cumulative parameter updates and projects gradients and optimizer steps to limit cross-task interference in multi-teacher on-policy distillation. A lightweight conflict probe guides task ordering, while a cycling strategy balances subspace estimation with revisiting tasks. On Code, Reason, and Math, PMOPD raised the three-task average by 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B versus standard multi-teacher on-policy distillation.
- Task updates concentrate in separate low-dimensional subspaces during multi-teacher distillation.
- PMOPD projects gradients and optimizer steps away from protected task directions.
- A conflict probe and cycling strategy order tasks and revisit them.
- Average score rose 2.54 points on Qwen2.5-7B and 2.09 on Llama-3.1-8B.
Full article214 words · extracted from huggingface.co · click to collapse
Multi-teacher on-policy distillation (MOPD) has emerged as a popular post-training paradigm for integrating specialized capabilities in frontier language models. Existing OPD research has primarily focused on optimizing single-task distillation through objective design, distillation scope, and teacher signal construction, whereas MOPD must aggregate multiple capabilities in shared parameters and address the resulting capability seesaw, in which improving one domain suppresses capabilities acquired from another. Inspired by the distinctive update geometry of OPD, we find that parameter updates from different tasks rapidly concentrate in their respective low-dimensional subspaces during MOPD, providing a direct geometric basis for identifying and controlling cross-task interference. We therefore propose PMOPD (Projection-based Multi-Teacher On-Policy Distillation), which constructs subspace memories from the cumulative parameter displacements of different tasks and projects both gradients and optimizer updates to remove components that interfere with protected task directions. We further develop a lightweight conflict probe to characterize task interactions and guide task ordering, together with a cycling strategy that balances subspace estimation and timely task revisitation. Experiments on representative Code, Reason, and Math tasks show that PMOPD improves every evaluated capability over MOPD, raising the average score across the three tasks by 2.54 points on Qwen2.5-7B and 2.09 points on Llama-3.1-8B. These consistent gains establish geometry-aware optimization as an effective and transferable approach to balanced multi-teacher distillation.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.34605