Latent-MOPD: Latent Multi-Teacher On-Policy Distillation
Latent-MOPD distills multiple LLM specialists via hidden states and token predictions, beating baselines on nine benchmarks.
Latent-MOPD is presented as the first representation-level multi-teacher on-policy distillation method for LLMs, combining specialists' output distributions and hidden states without extra teacher training. Late-layer targets, a shared projection for unequal widths, and domain-grouped updates coordinate supervision, which gradually shifts from hidden states to token predictions. In same-family tests it beat token-only, representation-only, and uniform-averaging baselines on all nine math, code, and logic benchmarks, and with equal parameters surpassed the best teacher on most of them. With larger cross-family teachers it also beat both single-channel baselines on every benchmark.
- First representation-level multi-teacher on-policy distillation method for LLMs.
- Uses late-layer hidden states and token predictions from one routed specialist.
- Beats token-only, representation-only, and averaging baselines on all nine benchmarks.
- Equal-parameter student surpasses the best teacher on a majority of benchmarks.
- An all-layer representation-only control collapses when teacher domains are interleaved.
Full article204 words · extracted from huggingface.co · click to collapse
On-policy distillation (OPD) trains a student on the responses it generates. Existing LLM multi-teacher OPD transfers what specialists predict through their output distributions. We introduce Latent-MOPD, to our knowledge the first representation-level multi-teacher OPD method for LLMs. It integrates existing specialists through both their predictions and the hidden states used to compute them, without additional teacher training. To coordinate representation supervision from multiple specialists, we select late-layer targets according to the teacher-student relationship, bridge unequal hidden widths with a shared projection, and group updates by domain. Each teacher's supervision gradually shifts from hidden states to token predictions, with both channels using the same routed specialist. In our main same-family setting, Latent-MOPD outperforms the token-only, representation-only and uniform-averaging baselines on all nine benchmarks across math, code and logic. With the same parameter count as each teacher, the student also surpasses the per-benchmark best teacher on a majority of these benchmarks. With larger, separately developed cross-family teachers, Latent-MOPD outperforms both single-channel baselines on all benchmarks. A same-family all-layer representation-only control remains stable with domain-pure updates but collapses when teacher domains are interleaved within an update. Our results show that a single student can integrate capabilities from several specialists through both their output distributions and internal representations.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.02381