Post-Training Leaves Behavioral Shadows on Unrelated Decisions
Active Taskless Distillation transfers coding and reasoning skills using one unrelated teacher word per prompt.
The paper introduces Active Taskless Distillation, which transfers post-training capabilities using a single ordinary word from the teacher on prompts where a shared ancestor model is nearly indifferent. A student learns only from those prompt-word pairs, without target-task examples, teacher logits, or teacher parameters. In the main coding experiment, Qwen2.5-1.5B gains 5.34 percentage points on HumanEval+ over a nuisance-matched control. Further experiments report transfer in scientific knowledge, commonsense reasoning, and reading comprehension across model generations, sizes, and families, with a composable signal that tracks teacher update strength.