Learning from Teacher Continuations at Student States
OLIVE updates a student on teacher continuations of student-generated prefixes, beating offline distillation under the same budget.
Researchers present OLIVE, an online distillation method where the student generates a prefix, a teacher continues it autoregressively, and the student is trained with cross-entropy on teacher tokens. OLIVE targets covariate shift in offline SFT, fragmented supervision in on-policy distillation, and the need for teacher token probabilities. It outperforms on-policy distillation at comparable GPU-hour cost, and an asynchronous implementation cuts total training time by 23.8%. Using only text from GPT-5.4-mini, continuous OLIVE training beats offline SFT from the same teacher by 13% on ScienceWorld.
- Student prefixes are continued by the teacher and trained with cross-entropy.
- Asynchronous OLIVE cuts total training time by 23.8 percent.
- GPT-5.4-mini text plus OLIVE beats offline SFT by 13% on ScienceWorld.
- OLIVE keeps improving after offline distillation plateaus.
Full article195 words · extracted from huggingface.co · click to collapse
We present OLIVE (OnLine InterVEntion). At each iteration, the evolving student policy generates a new prefix, the teacher continues it autoregressively, and the student is updated using cross-entropy computed on the teacher-generated tokens. Each design choice targets a corresponding limitation of existing distillation methods: (1) sequential covariate shift in offline supervised fine-tuning (SFT) on fixed teacher trajectories, (2) fragmented supervision under prefix failure in token-level on-policy distillation (OPD), and (3) the need for access to teacher token probabilities in distribution-matching distillation. OLIVE achieves higher reasoning performance than OPD (with a top-16 KL approximation) at comparable GPU-hour cost. Our asynchronous implementation further reduces OLIVE's total training time by 23.8\%. We evaluate OLIVE on both hard reasoning tasks and agentic tasks which reflects modern post-training scenarios, and it consistently outperforms existing distillation methods under the same training budget. By regenerating prefixes from the evolving student, OLIVE continues improving after offline distillation plateaus while better preserving the general capabilities and plasticity of the student. Using only text from GPT-5.4-mini, continuously training with OLIVE outperforms offline SFT from the same teacher by 13\% on ScienceWorld. These results support OLIVE as an effective and efficient approach to online language-model distillation.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.36246