Composing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts
Δ-MOPD distills multiple teachers by transferring teacher-minus-base logit shifts re-anchored at the student's frozen initialization, beating endpoint-policy composition by 4.11 Math points across three teachers.
The paper introduces Δ-MOPD, a multi-teacher on-policy distillation method. Rather than copying each teacher's endpoint policy, it transfers the teacher-minus-base logit shift re-anchored at the student's frozen initialization. The authors identify inherited base pull (inherited base preferences) as the mechanism that impedes endpoint transfer, showing it can exceed the teacher's post-training shift. With three composed teachers, Δ-MOPD exceeds endpoint composition by 4.11 Math points and 1.95 points across five benchmarks; with two teachers it matches endpoint accuracy. Phased routing raises mean performance and cuts the teacher order gap from 10.50 to 6.42 points, while interleaved routing shows comparable results for both transfer targets. The story appeared on Hugging Face daily papers on 2026-10-06 and on arXiv (cs.AI / cs.LG / cs.CL) on 2026-10-07; the two reports agree on all figures.
- Method: Δ-MOPD (multi-teacher on-policy distillation) transfers teacher-minus-base logit shifts re-anchored at the student's frozen initialization, not endpoint policies
- Mechanism: inherited base pull can exceed the teacher's post-training shift, impeding endpoint transfer
- Three-teacher composition gains 4.11 Math points and 1.95 points across five benchmarks over endpoint composition
- With two teachers, Δ-MOPD matches endpoint accuracy
- Phased routing reduces the teacher order gap from 10.50 to 6.42 points
- Interleaved single-teacher updates perform similarly for both transfer targets
- Sources: Hugging Face daily papers (2026-10-06) and arXiv cs.AI/cs.LG/cs.CL (2026-10-07); no disagreements between reports
Coverage timelineoldest first · each row is one article
- · 2d agoComposing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts
Hugging Face daily papers· 32
Δ-MOPD distills multiple teachers via teacher-relative logit shifts instead of endpoint policies, beating endpoint composition by 4.11 Math points.
- · 1d agoComposing What Each Teacher Learned: Multi-Teacher On-Policy Distillation through Teacher-Relative Shifts
arXiv cs.AI / cs.LG / cs.CL· 44
Δ-MOPD transfers teacher-minus-base logit shifts, beating endpoint multi-teacher distillation on math and mixed benchmarks.