Beyond Teacher Assignment: Domain-Normalized Multi-Teacher On-Policy Distillation
Domain-normalized multi-teacher distillation rebalances specialist feedback so Qwen3.5 students retain mathematics gains.
The paper studies multi-teacher on-policy distillation, where domain specialists teach one student model. On Qwen3.5 at three sizes, unbalanced instruction-following feedback dominated updates, so the student did not beat the best single specialist and lost most of the mathematics gain. Domain-Normalized MOPD keeps the same routing but rescales each domain's feedback by its measured spread, improving average scores on six public benchmarks across three seeds and two answer-length limits. Controls indicate the gain comes mainly from reducing instruction-following feedback rather than boosting mathematics alone.
- MOPD students failed to beat the best single specialist
- Instruction-following feedback was far more dispersed than math feedback
- DN-MOPD rescales each domain's feedback by its measured spread
- Average scores rose on six benchmarks at every Qwen3.5 size
- Fixed weights near the measured scales performed comparably
Full article227 words · extracted from huggingface.co · click to collapse
Reinforcement learning can turn one language model into several specialists, each excellent at a single skill such as mathematics, coding or following instructions, but users need one model with all of these skills. Multi-teacher on-policy distillation (MOPD) merges them by letting the specialists teach one student: the student answers each prompt, and the specialist for that prompt's domain gives feedback on every token. This routing decides which specialist teaches, but not how strongly its feedback moves the shared student. In Qwen3.5 models at three sizes, we find that MOPD's student does not beat one taught by the best single specialist and gains little of the mathematics specialist's advantage. The feedback is unbalanced: instruction-following feedback is several times more spread out than mathematics feedback and dominates the student's updates. We propose Domain-Normalized MOPD (DN-MOPD), which keeps the routing and rescales each domain's feedback by its measured spread. On six public benchmarks, DN-MOPD improves the average score over MOPD at every size, across three random seeds and under two answer-length limits, and recovers most of the lost mathematics gain. Controls with fixed domain weights show that the gain comes mainly from turning down instruction-following feedback rather than turning up mathematics alone, and that fixed weights close to those DN-MOPD measures perform comparably. Combining specialists therefore requires deciding not only which one teaches, but also how strongly its feedback counts.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.35347