FuseReg: Regularizing Layer Fusion Mitigates the Reconstruction-Generation Gap in Representation Autoencoders
FuseReg trains decoders on random encoder-layer subsets, narrowing the reconstruction-generation gap in RAEs.
FuseReg addresses the trade-off in representation autoencoders between shallow encoder layers, which preserve pixels, and deeper layers, which improve generation. It trains on random subsets of encoder layers so one decoder can reconstruct from full, sparse, and single-layer fusions. On ImageNet-256 with DINOv3-L, it reports higher PSNR than decoders specialized to fixed fusions. Replacing only the decoder reduces unguided gFID by 27% for an unchanged RAEv2 DiT-XL generator, and joint regularization reduces it by 29% on DiT-Base.
- Random layer-subset training reduces sensitivity to cross-layer disagreement.
- One decoder handles full, sparse, and single-layer fusions without retraining.
- Decoder replacement cuts unguided gFID by 27% on RAEv2 DiT-XL.
- Joint regularization cuts unguided gFID by 29% on DiT-Base.
Full article194 words · extracted from huggingface.co · click to collapse
Representation autoencoders (RAEs) reuse features from a pretrained visual encoder as reconstruction and diffusion latents, integrating strong visual representations into image generation. However, RAEs still need to decide which encoder layers form the shared latent space for the generator and pixel decoder. This choice involves a trade-off. Shallower layers tend to preserve fine pixel details better, while deeper layers tend to yield better generation metrics. A fixed heuristic layer fusion therefore couples two stages that benefit from different information. We introduce FuseReg, which replaces heuristic feature selection with training over random subsets of encoder layers. We theoretically analyze the underlying mechanism: subset sampling explicitly penalizes sensitivity to cross-layer disagreement. On ImageNet-256 with DINOv3-L, a single FuseReg decoder reconstructs from full, sparse, and single-layer fusions without retraining, achieving higher PSNR than decoders specialized to fixed fusions. This flexibility also benefits generation: decoder replacement alone reduces unguided gFID by 27% with an unchanged RAEv2 DiT-XL generator. The same regularization principle extends to diffusion training, with joint regularization of both stages reducing unguided gFID by 29% on DiT-Base. These results show that training downstream models for layer-fusion robustness narrows the reconstruction-generation gap without modifying the pretrained encoder.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.31620