Distribution Matching Distillation for Continuous Diffusion Language Models
Simplex-DMD distills continuous diffusion language models to 4 evaluations with perplexity 45.6.
The paper distills continuous diffusion language models, which can need hundreds of network evaluations, using two reverse-KL methods with the same student architecture. Simplex-DMD uses continuous token relaxations and pathwise gradients; Reinforce-DMD uses categorical sampling and REINFORCE with a learned density ratio. On OpenWebText sequences of 1,024 tokens, Simplex-DMD reaches generative perplexity 45.6 at 5.44 nats entropy in 4 NFEs, a 49% reduction versus the strongest matched diffusion baseline. Reinforce-DMD reaches perplexity 14.9 at 5.00 nats with 256 NFEs, a 20% reduction.
- Simplex-DMD and Reinforce-DMD share reverse-KL distillation
- Simplex-DMD reaches perplexity 45.6 in only 4 NFEs
- That is a 49% perplexity cut versus the diffusion baseline
- Reinforce-DMD reaches perplexity 14.9 using 256 NFEs
Full article171 words · extracted from arxiv.org · click to collapse
Continuous diffusion language models generate all tokens in parallel, yet high-quality generation can still require hundreds of network evaluations (NFEs). We study how distributional distillation can reduce this cost by exploiting the student's probabilistic token outputs. Our unified formulation connects the student's output parameterization to the resulting gradient estimators and yields two methods with the same student architecture and reverse-KL matching objective: Simplex-DMD uses continuous token relaxations and pathwise gradients, while Reinforce-DMD uses categorical sampling and REINFORCE with a learned density ratio. We develop both methods for multi-step generation and investigate the training and sampling choices associated with each parameterization. On OpenWebText, for sequences of 1,024 tokens, Simplex-DMD achieves a generative perplexity of 45.6 at a unigram entropy of 5.44 nats in just 4 NFEs, a 49% reduction relative to the strongest evaluated diffusion baseline at matched entropy and sampling budget. Reinforce-DMD improves the frontier at larger budgets, reaching a generative perplexity of 14.9 at an entropy of 5.00 nats with 256 NFEs, a 20% reduction under the same comparison protocol.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.40235