DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation
DMAD distills few-step image and video generators via adversarial distribution matching, reporting top compared scores.
DMAD recasts distribution-matching distillation as adversarial classification so a few-step student can be trained without fitting an auxiliary diffusion model to the student distribution. Two discriminator heads estimate the needed log-density ratios, and gap-based reweighting adjusts teacher supervision across noise levels. Reported results include FID 1.04 for one-step ImageNet-64x64 generation, FID 14.47 for four-step SDXL on COCO-10K, and a VBench total of 85.15 for four-step Wan2.1-T2V-14B. On MiniMax-H3-33B, the four-step student wins 79.1% of comparisons against DMD2 and 84.6% against rCM for joint audio-video generation, excluding ties.
- DMAD learns distribution-matching gradients with discriminator heads, without an auxiliary score model.
- One-step ImageNet-64x64 generation reaches FID 1.04.
- Four-step SDXL scores FID 14.47 on COCO-10K; four-step Wan2.1 scores 85.15 on VBench.
- A four-step MiniMax-H3-33B student is preferred over DMD2 and rCM in human ratings.
Full article212 words · extracted from huggingface.co · click to collapse
Distribution Matching Distillation (DMD) trains a few-step student from the difference between separately estimated target and student scores, so it must keep an auxiliary diffusion model fitted to the student's evolving distribution at extra memory and computation cost. We introduce DMAD, Distribution Matching as Adversarial Distillation, which recasts distribution matching as classification and learns the required log-density ratios directly. Two discriminator heads on a shared backbone distinguish real data and teacher samples from the student's, and linear losses on their logits train the student without auxiliary score fitting. We prove that at the discriminator optimum these losses recover the distribution-matching gradient underlying DMD, through the classical identity linking discriminator logits to log-density ratios. We further introduce gap-based reweighting, which adapts teacher supervision across noise levels from the real-data head's empirical logit gap between real and teacher samples. DMAD reaches a Fréchet Inception Distance (FID) of 1.04 with one-step generation on ImageNet-64x64, 14.47 with four-step SDXL on COCO-10K, and a VBench total score of 85.15 with four-step Wan2.1-T2V-14B, the best values among the compared few-step methods and the multi-step teachers. On MiniMax-H3-33B, our four-step student achieves overall human preference rates of 79.1% over DMD2 and 84.6% over rCM for joint audio-video generation, excluding ties. Our code, models and demos are available at https://yzmblog.github.io/projects/DMAD.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.02188