Post-Training Leaves Behavioral Shadows on Unrelated Decisions
Active Taskless Distillation transfers coding and reasoning skills using one unrelated teacher word per prompt.
The paper introduces Active Taskless Distillation, which transfers post-training capabilities using a single ordinary word from the teacher on prompts where a shared ancestor model is nearly indifferent. A student learns only from those prompt-word pairs, without target-task examples, teacher logits, or teacher parameters. In the main coding experiment, Qwen2.5-1.5B gains 5.34 percentage points on HumanEval+ over a nuisance-matched control. Further experiments report transfer in scientific knowledge, commonsense reasoning, and reading comprehension across model generations, sizes, and families, with a composable signal that tracks teacher update strength.
- ATD uses one ordinary teacher word per prompt, with no task examples.
- Students see neither teacher logits nor teacher parameters.
- Qwen2.5-1.5B gained 5.34 points on HumanEval+ versus a matched control.
- Transfer also appears in knowledge, commonsense, and reading tasks.
- The learned signal is composable and tracks the teacher's update strength.
Full article167 words · extracted from huggingface.co · click to collapse
We find that language models can transfer capabilities through task-unrelated text. Post-training typically improves language models using task-specific data. Prior work on subliminal learning shows that information about these updates can pass through unrelated generations, but has largely focused on traits or preferences using extensive teacher outputs. We introduce Active Taskless Distillation (ATD), which achieves capability transfer using only a single word from the teacher per prompt. ATD probes the behavioral shadow of post-training by selecting prompts where the teacher and student's shared public ancestor is nearly indifferent between two ordinary words. A student initialized from this ancestor learns solely from the resulting prompt-word pairs, without target-task examples, teacher logits, or teacher parameters. In the primary coding experiment with Qwen2.5-1.5B, 5,664nses yield a 5.34 pp gain on HumanEval+ over an exact nuisance-matched control thadisrupts prompt-resperiments showtransfer in scientific knowledge, commonsense reasoning, and reading comprehensins across additional model generations, sizes, and families. Functional analyses show that the learned sid composable, andthat its strength tracks the teacher's update strength.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.29233