Embedding Prediction Helps Image Generation
NEPA-DiT-XL predicts embeddings each denoising step and reaches ImageNet FID 1.32 with about one-third of REPA's compute.
Next-Embedding Predictive Autoregression trains a Transformer to predict the next continuous embedding rather than reuse one fixed class or text condition. In Embedding Conditioned Generation, a diffusion transformer is conditioned on Multi-Embedding Prediction outputs that are recomputed at every denoising step. On class-conditional ImageNet 256x256, NEPA-DiT-XL combined with REPA reaches an FID of 1.32 while using about one-third of REPA's training compute, at the cost of a second network on each sampling step.
- NEPA predicts next continuous embeddings instead of reusing a fixed class or text condition.
- Multi-Embedding Prediction forecasts clean-image embeddings from the condition and noisy image.
- The diffusion generator is reconditioned on those predictions at every denoising step.
- NEPA-DiT-XL reaches FID 1.32 on ImageNet 256 using about one-third of REPA compute.
Full article173 words · extracted from arxiv.org · click to collapse
In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to predict them all at once with Multi-Embedding Prediction, and in Embedding Conditioned Generation, a DiT generator is conditioned on these predictions, recomputed at every denoising step, so the conditioning signal adapts to the current noisy state. Experiments on class-conditional ImageNet $256\times256$ study the condition of the generator, the design of Multi-Embedding Prediction, and the scaling of both models. The NEPA model adds a second network to every sampling step; with it, and combined with REPA, our final model, NEPA-DiT-XL, reaches an FID of 1.32 using about a third of the training compute of REPA.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.02203