Enhancing Diffusion Language Models with Autoregressive Post-Training Weights
A2D recycles autoregressive post-training weights to improve diffusion language models without extra training.
The paper shows that autoregressive post-training weight updates can be added directly to diffusion language models after AR-to-diffusion conversion, approaching the gains of diffusion-native post-training. AR and diffusion updates are nearly orthogonal in weight space but complementary, so composing them retains both. The training-free A2D method improves instruction following, mathematical reasoning, and coding on Dream, DreamReasoner, DiffuCoder, Dream-Coder, Nemotron-Labs-Diffusion, and DiffusionGemma using existing SFT and RL updates, with no added training or inference cost.