Improved Distributional Diffusion Models
Improved distributional diffusion models reach 2.38 FID on ImageNet-256 with one from-scratch DiT.
The paper makes Distributional Diffusion Models practical by deferring particle expansion to late transformer layers and using time-dependent scoring-rule schedules. A single from-scratch DiT-XL/2 latent model reaches 4.48 FID at 4 steps and 2.38 FID at 50 steps on class-conditional ImageNet-256, without a teacher, self-distillation, or Jacobian-vector products. FID does not degrade as the sampling budget grows from 4 to 50 evaluations, and the recipe also transfers to text-to-image generation. Code and pre-trained models are released by CompVis.
- Late-layer particle expansion reduces multi-particle training overhead.
- Time-dependent scoring schedules replace one fixed hyperparameter trade-off.
- DiT-XL/2 reaches 4.48 FID at 4 steps and 2.38 at 50.
- The model is trained from scratch in one stage without a teacher.
Full article178 words · extracted from huggingface.co · click to collapse
Distributional Diffusion Models (DDMs) replace the standard mean-prediction denoiser with a distributional denoiser trained via a scoring rule objective, learning a stochastic approximation to p(x_1 mid x_t) rather than its conditional mean. However, scaling DDMs to modern image-generation settings faces two obstacles: (i) multi-particle training incurs overhead that scales with the number of particles, (ii) DDMs use globally fixed scoring rule hyperparameters, forcing a single trade-off across sampling budgets. We mitigate these limitations by deferring particle expansion to late transformer layers, and the hyperparameter trade-off by introducing time-dependent scoring rule schedules informed by the dynamical regimes of~Biroli2024. Combined with a DiT-based latent setup, these changes make DDM training practical on class-conditional ImageNet-256^2, achieving 4.48 FID at 4 steps and 2.38 at 50 steps with DiT-XL/2, from a single model trained from scratch in one stage, without a teacher, self-distillation or JVPs. The result is a stochastic few-step generator whose FID does not degrade as the sampling budget grows from 4 to 50 NFE, and the same recipe transfers to text-to-image generation. Code and pre-trained models available at https://github.com/CompVis/iDDM.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.37147