Pivot-SD: Efficient Self-Distillation for Masked Diffusion Language Models
Pivot-SD self-distills masked diffusion models by training only high-impact commitments, improving LLaDA-8B on math and code.
Pivot-SD is an offline self-distillation method for masked diffusion language models that trains only high-impact token commitments, or pivots, rather than whole sequences or denoising steps. Pivots are selected with an information-gain metric that measures how much a commitment reduces uncertainty over remaining masked positions. Successful-trajectory pivots are trained with cross-entropy and failed-trajectory pivots with targeted unlikelihood. Using 200 questions and four rollouts each, it improved LLaDA-8B-Instruct over full-sequence SFT and budget-matched diffusion RL on math and code benchmarks.
- Pivot-SD supervises only high-impact denoising commitments called pivots.
- Pivots are chosen by an information-gain uncertainty-reduction metric.
- Successful pivots use cross-entropy; failed ones use targeted unlikelihood.
- Training used 200 questions with four rollouts each.
- LLaDA-8B-Instruct beat full-sequence SFT and budget-matched diffusion RL.
Full article163 words · extracted from huggingface.co · click to collapse
Masked diffusion language models (dLMs) offer a promising parallel alternative to autoregressive models for complex reasoning. However, they face a distinct credit-assignment challenge, since a few commitments during denoising sharply reduce the uncertainty over the remaining masked positions and shape much of the response. Most post-training recipes for dLMs do not use this signal to decide which tokens to train on: they typically train on the final text or assign rewards to whole denoising steps, rather than selecting the individual commitments that shape the response. We introduce Pivot-SD, an efficient offline self-distillation framework that supervises only these high-impact commitments (pivots). Pivot-SD selects pivots using an information-gain metric measuring uncertainty reduction over the remaining masked positions. Pivots from successful trajectories are trained with cross-entropy, and pivots from failed trajectories with targeted unlikelihood, leaving the rest of the failed trajectory untouched. Using only 200 questions and four rollouts each, Pivot-SD improves LLaDA-8B-Instruct over full-sequence SFT and budget-matched diffusion RL baselines across math and code benchmarks.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.03665