Representation-Space MMD for Diffusion Language Models
Representation-space MMD post-training improves diffusion language models without full sampling trajectories.
The paper proposes post-training for diffusion language models that minimizes Maximum Mean Discrepancy between generated and reference distributions in a frozen model's feature space. Token-level contextual features provide multiple observations per sequence from one extractor pass, using policy gradients for discrete models and direct differentiation for continuous models, without full sampling trajectories or auxiliary models. Experiments show lower generative perplexity at comparable entropy on OpenWebText and better accuracy-computation trade-offs on GSM8K. On 16B DMax-LLaDA2.0 models with hybrid masked-uniform diffusion, higher decoding parallelism maintained similar or higher math and code accuracy.
- MMD is estimated from token features of a frozen diffusion language model.
- Discrete models use policy gradients; continuous models differentiate through latents.
- OpenWebText perplexity falls at similar entropy; GSM8K accuracy-compute trade-offs improve.
- 16B DMax-LLaDA2.0 keeps or raises math and code accuracy with more parallelism.
Full article130 words · extracted from huggingface.co · click to collapse
We introduce a post-training method for diffusion language models (DLMs) that minimizes Maximum Mean Discrepancy (MMD) between generated and reference distributions in the feature space of a frozen pretrained DLM. To estimate MMD, we retain contextual features at individual token positions, obtaining multiple observations per sequence from a single extractor pass. We optimize this objective using policy gradients for discrete models and direct differentiation through generated latents for continuous models. In both cases, computing the loss directly from these features enables efficient post-training without full sampling trajectories or jointly trained auxiliary models. Experiments show lower generative perplexity at comparable entropy on OpenWebText and better accuracy-computation trade-offs on GSM8K. On 16B DMax-LLaDA2.0 models with hybrid masked-uniform diffusion, we increase decoding parallelism with similar or higher accuracy on math and code benchmarks.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.06648