GRACE: Generation-aware latent compression for efficient video generation
GRACE cuts Wan2.1-I2V-14B video tokens 8x and latency 11.1x without losing VBench quality.
GRACE compresses a pretrained video autoencoder while keeping it compatible with a pretrained Diffusion Transformer. It freezes the base latent, learns a residual for information lost under stronger compression, and aligns the compressed latent with the original in the frozen DiT feature space. The DiT is then adapted with lightweight fine-tuning and asymmetric denoising, denoising the base ahead of the residual. On Wan2.1-I2V-14B, GRACE reduces token count 8x and latency 11.1x at 480x832x81 while matching the pretrained pipeline's VBench quality.
- GRACE keeps a frozen base latent and learns a residual for lost detail.
- Latents are aligned in the frozen DiT feature space for generation quality.
- Wan2.1-I2V-14B token count drops 8x and latency 11.1x at 480x832x81.
- Lightweight fine-tuning and asymmetric denoising match pretrained VBench quality.
Full article239 words · extracted from huggingface.co · click to collapse
Highly compressed video autoencoders offer an effective way to accelerate video diffusion models, as the Diffusion Transformer (DiT) operates on far fewer tokens. However, such autoencoders are challenging to train, since a higher compression ratio degrades reconstruction quality and recovering it requires more channels, which is known to slow the convergence of the DiT. The compressed latent also differs from the one the DiT was trained on, so the pretrained DiT must be either retrained from scratch or adapted at considerable cost. Compressing the autoencoder the DiT was trained with appears to preserve compatibility, yet optimizing it for reconstruction alone still shifts the latent away from the distribution the DiT has learned. To address this, we propose Generation-Aware Latent Compression for Efficient Video Generation (GRACE), a two-stage framework that compresses a pretrained video autoencoder while keeping it compatible with the pretrained DiT. Specifically, we keep a frozen base latent from the pretrained encoder and learn a residual latent for the information lost under stronger compression, while aligning the compressed latent with the pretrained latent in the feature space of the frozen DiT so that the autoencoder is optimized for generation. We then adapt the DiT with lightweight fine-tuning and asymmetric denoising, where the base is denoised ahead of the residual. GRACE reduces the token count of Wan2.1-I2V-14B by 8x and its latency by 11.1x at 480x832x81, while matching the generation quality of the pretrained pipeline before compression on VBench.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.10524