MOPD-Router: Rethinking Teacher Routing in Multi-Teacher On-Policy Distillation
MOPD-Router routes multi-teacher distillation per token without domain labels, and ExpertAlign beats prior routing.
MOPD-Router routes on-policy distillation signals across a full pool of teachers at each token, without prompt-level domain labels or a separately trained router. Its ExpertAlign metric scores whether a teacher's correction expresses that teacher's post-training specialization, and it is compared with entropy and novelty baselines. ExpertAlign is strongest in all four strong-to-weak and same-size settings: plus 5.88 points (+12.3%) over mean aggregation on unlabeled data, and plus 3.95 points (+7.8%) over standard MOPD on domain-labeled data without using those labels. Code is published at github.com/TURLEing/MOPD-Router.
- MOPD-Router assigns teacher supervision at every token without domain labels.
- ExpertAlign scores whether a teacher's correction matches its specialization.
- It beats mean aggregation by 5.88 points on unlabeled mixtures.
- On labeled data it beats standard MOPD by 3.95 points without using labels.
Full article207 words · extracted from huggingface.co · click to collapse
Multi-teacher on-policy distillation (MOPD) integrates specialized capabilities into a single student, but existing practice typically hard-routes each prompt to a domain-matched teacher for the entire rollout. This dependence on prompt-level domain labels restricts using unlabeled training mixtures and leaves complementary signals from other teachers unused. We introduce MOPD-Router, a framework that routes supervision over the full teacher pool at each token, without domain labels or training a separate routing model. Its plug-in interface supports different metrics for selecting and weighting teacher-specific OPD signals. Within this interface, we propose ExpertAlign, which scores each teacher by whether its correction to the student at the current token expresses the specialization that teacher acquired during post-training, and compare it against two reference metrics built on teacher confidence (Entropy) and teacher-student discrepancy (Novelty). Experiments on unlabeled and domain-labeled training mixtures under strong-to-weak and same-size distillation scenarios show that ExpertAlign achieves the strongest overall performance in all four settings. On unlabeled data, it improves the overall score by 5.88 (+12.3%) points over Mean aggregation; on domain-labeled data, it outperforms standard MOPD by 3.95 (+7.8%) points without using available domain labels. These results demonstrate token-level routing can exploit cross-domain complementary supervision, and reduce exclusive reliance on prompt-level domain assignment. Code is available at: https://github.com/TURLEing/MOPD-Router.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.30837