From Gradients to Capabilities: Understanding Multi-Teacher On-Policy Distillation
A study of Qwen3-1.7B shows multi-teacher on-policy distillation is shaped by averaging, Adam, and BF16 precision.
The paper analyzes multi-teacher on-policy distillation, which combines RL-trained teachers into one student. Using Qwen3-1.7B and four domain teachers from the same initialization, plus SmolLM3-3B diagnostics, the authors compare gradients, optimizer updates, and learning curves. Loss averaging weights longer responses more heavily, while Adam's first moment raises cosine similarity of parameter updates from 0.83 between teachers to 0.96 between averaging rules. BF16 hides most change—about 97% of FP32 master weights differ from initialization versus 7–11% of BF16 weights—and math accuracy shifts by 2.6 or 2.1 points depending on averaging.
- Qwen3-1.7B student is compared with four RL domain teachers from the same initialization.
- Token averaging favors longer responses; Adam update similarity reaches 0.96 across averaging rules.
- About 97% of FP32 master weights change, but only 7-11% of BF16 weights.
- Top-64 KL matches the full-vocabulary gradient; math accuracy shifts by about two points.
Full article172 words · extracted from arxiv.org · click to collapse
Multi-teacher on-policy distillation (MOPD) aims to combine the strengths of RL-trained teachers in a single student, but how teacher signals affect parameter changes remains underexplored. We study Qwen3-1.7B with four domain teachers trained with RL from the same initialization as the student, comparing gradients, optimizer updates, and task learning curves, with additional SmolLM3-3B diagnostics. We find that several factors influence teacher signals. First, loss averaging implicitly weights responses: token averaging favors longer responses, and equalizing domain contributions retains this weighting within domains. Second, Adam's first moment reduces differences in parameter updates: the cosine similarity is 0.83 between teachers and 0.96 between averaging rules, despite differences in raw gradients. Third, BF16 rounding hides small changes: about 97\% of FP32 master weights differ from initialization, but only 7--11\% of BF16 weights do. Finally, the top-64 intersection KL gradient closely matches Qwen's full-vocabulary gradient, but the effect on task performance depends on averaging: mathematics accuracy is 2.6 points higher than with sampled-token policy-gradient (PG) under response averaging and 2.1 points lower under global token averaging.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.02179