DMAD: Distribution Matching as Adversarial Distillation for Fast Visual Generation
DMAD distills few-step image and video generators by learning distribution matching with adversarial discriminators.
DMAD turns distribution-matching distillation into classification, training a few-step student from discriminator logit losses without fitting an auxiliary score model. Two heads on a shared backbone separate real and teacher samples from student samples, and gap-based reweighting adapts teacher supervision across noise levels. Reported scores include FID 1.04 for one-step ImageNet-64, FID 14.47 for four-step SDXL on COCO-10K, and VBench 85.15 for four-step Wan2.1-T2V-14B. A four-step MiniMax-H3-33B student wins 79.1% human preference over DMD2 and 84.6% over rCM on joint audio-video, excluding ties.