Unifying Distributional Training for One-Step Visual Generation
MGFlow unifies one-step visual training and beats FD-Loss on ImageNet and four-step FLUX.2.
The paper presents a unified distributional-training framework for one-step visual generation that separates distribution modeling from matching discrepancy using Wasserstein gradient flow. Under the framework, FD-Loss and Gaussian-kernel Drifting emerge as special cases, motivating MGFlow, which models features with Gaussian mixtures and pairs mass-constrained assignment with component updates to limit mode collapse. On ImageNet 256x256, MGFlow reports 1.45 FDr^6 on pMF-H and 1.64 on JiT-H, surpassing the FD-Loss baseline. Post-training FLUX.2 [klein] 4B into a one-step generator outperforms the original four-step model on GenEval and PickScore.
- Framework recovers FD-Loss and Gaussian-kernel Drifting as special cases.
- MGFlow uses Gaussian mixtures with mass-constrained sample assignment.
- ImageNet 256: 1.45 FDr^6 on pMF-H and 1.64 on JiT-H.
- One-step FLUX.2 klein 4B beats the four-step model on GenEval and PickScore.
Full article158 words · extracted from arxiv.org · click to collapse
\emph{Distributional training} provides collective supervision for one-step visual generation by matching real and generated features in frozen representation spaces. We introduce \emph{a unified theoretical framework} that separates distribution modeling from matching discrepancy and connects global objectives to pointwise feature updates through Wasserstein gradient flow. Under this framework, FD-Loss and Gaussian-kernel Drifting are recovered through Gaussian optimal transport and kernel-density-based KL matching, respectively. The framework motivates \textbf{MGFlow}, which models feature distributions with Gaussian mixtures at an adjustable granularity between global moments and sample-based representations. MGFlow supports both optimal transport and score-based matching, and couples mass-constrained sample assignment with paired component updates to address mode collapse that mixture expressivity alone does not resolve. On ImageNet $256\times256$, MGFlow substantially surpasses the FD-Loss baseline, achieving state-of-the-art results with \textbf{1.45} $\mathrm{FDr}^6$ on pMF-H and \textbf{1.64} on JiT-H. For text-to-image generation, MGFlow post-trains FLUX.2 [klein] 4B into a one-step generator that outperforms the original four-step model on both GenEval and PickScore. Project page: https://shihaoyang0423.github.io/MGFlow-website/
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.35763