Persistence Forcing: Exploiting Feature Specialization in Pixel-Space Diffusion
Persistence Forcing uses specialized diffusion features to reach FID 1.63 on ImageNet 256.
Pixel-space diffusion Transformers typically refine features uniformly, despite images needing compact global structure and richer local detail. Heterogeneous refinement produces persistent features for global structure and active features for high-frequency detail. Persistence Forcing lets persistent features condition active ones and adds guidance that complements classifier-free guidance. On ImageNet, PerF-L reaches FID 1.91 versus JiT-H's 1.86 with half the parameters, and PerF-H reaches 1.63 at 256×256 and 1.76 at 512×512.
- Heterogeneous depth budgets create persistent and active features
- Persistent features encode structure; active features encode detail
- PerF-H FID is 1.63 at 256 and 1.76 at 512
- PerF-L nears JiT-H quality with about half the parameters
Full article200 words · extracted from huggingface.co · click to collapse
Pixel-space diffusion Transformers (DiTs) directly operate on high-dimensional visual data, yet their hidden representations typically undergo uniform refinement across depth. Natural images, however, are inherently organized at different levels of granularity. Global structure can often be represented compactly, whereas local textures and fine details require richer representations. Motivated by this, we introduce heterogeneous refinement in pixel-space DiTs, assigning different feature groups distinct refinement budgets across depth. Consequently, an ordered feature specialization emerges: sparsely refined features predominantly encode global visual structure, whereas more frequently refined features increasingly specialize toward localized, high-frequency details. We refer to these two groups as persistent and active features, respectively. Building on this emergent specialization, we introduce Persistence Forcing (PerF), which explicitly exploits this persistent--active feature organization for pixel-space image generation. This enables persistent features to continuously condition actively refined features, allowing stable global information to guide the ongoing refinement of finer visual details. During generative sampling, this interaction further induces a meaningful guidance direction that promotes coherent global structure and naturally complements classifier-free guidance. On ImageNet 256times256, PerF-L achieves FID of 1.91, approaching 1.86 of JiT-H with only half the parameters, while PerF-H further achieves FID of 1.63 and 1.76 on ImageNet 256times256 and 512times512, respectively.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.36014