PDMD: Projected Distribution Matching Distillation for Video Diffusion Models
PDMD stabilizes video diffusion distillation by filtering critic errors, beating DMD at 4 NFE.
Projected Distribution Matching Distillation (PDMD) filters critic errors that degrade Distribution Matching Distillation for video diffusion by projecting out the update component parallel to the student-critic endpoint residual. The method needs only a one-line code change, with no extra loss, network, data, or multi-stage training. On Wan2.1, PDMD scores 83.73 VBench total at 4 NFE, 1.03 above matched DMD. On MiniMax-H3 joint video-audio generation it scores 83.17 VideoGen-Eval visual total and leads all six compared audio metrics among 4-NFE models.
- PDMD projects out critic-error components that accumulate in DMD student updates.
- Only a one-line DMD change; no extra loss, network, data, or stages.
- Wan2.1 scores VBench 83.73 at 4 NFE, 1.03 above matched DMD.
- MiniMax-H3 reaches VideoGen-Eval 83.17 and leads six audio metrics.
Full article241 words · extracted from arxiv.org · click to collapse
Modern video diffusion models require tens of denoising evaluations over long spatiotemporal token sequences. Distribution Matching Distillation (DMD) reduces the number of function evaluations (NFE) to just a few. However, DMD samples can degrade during training, exhibiting progressive oversaturation and artifacts. We trace this instability to critic errors, which enter successive student updates and accumulate over time. We introduce Projected Distribution Matching Distillation (PDMD) to filter critic errors. PDMD projects out the component of the DMD update parallel to the student-critic endpoint residual. At a fixed noisy query, we prove that this residual is an unbiased estimate of the critic's endpoint error. Under high-dimensional assumptions, this projection removes a constant fraction of critic error while discarding only a vanishing fraction of ideal DMD signal. Empirically, the projection stabilizes training and improves sample quality where DMD degrades and develops unnatural textures. PDMD requires only a one-line code change to DMD, with no extra loss, network, data, model pass, or multi-stage training. With Wan2.1, PDMD achieves a VBench total score of 83.73 at 4 NFE, surpassing matched DMD by 1.03 points. On MiniMax-H3 joint video-audio generation, PDMD achieves a VideoGen-Eval visual total score of 83.17, 0.41 points above the strongest distilled baseline. PDMD also achieves the best performance on all six audio metrics among the compared 4-NFE models. Qualitative comparisons and user studies favor PDMD over the distilled baselines in visual quality, motion, and audio quality. Code and models are available at https://pdmd2026.github.io/.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.35768