MC-Sparse: Deconstructing and Closing the Dense-Sparse Attention Gap in Diffusion Transformers
Training-free MC-Sparse speeds diffusion-transformer denoising up to 2.32× with negligible quality loss.
Meta-Cached Sparse Attention (MC-Sparse) is a training-free method that narrows the quality gap of sparse attention in diffusion transformers for video and high-resolution 3D generation. It selects individual key-value tokens, groups similar queries into tile-aligned blocks, and caches query groups, exact-probability KV indices, and dense-versus-sparse residuals for reuse across denoising steps. Relative to dense attention it reports a 1.80× denoising speedup on Minimax-H3-Base and 2.32× on 3D asset generation, with negligible quality loss and higher fidelity than prior sparse-attention baselines.
- Quality loss is traced to token grouping, bad selection, and discarded attention.
- MC-Sparse picks individual KV tokens and groups similar queries into GPU tiles.
- Cached groups, exact-probability KV indices, and residuals are reused across denoising steps.
- Reported speedups are 1.80× on Minimax-H3-Base and 2.32× for 3D generation.
Full article174 words · extracted from arxiv.org · click to collapse
Sparse attention is a primary approach to reducing the latency of diffusion transformers in long-sequence generation tasks, such as video and high-resolution 3D asset generation. However, existing methods can degrade generation quality and fidelity at high sparsity levels. Through controlled oracle comparisons, we trace this degradation to three sources: constraints imposed by token grouping, inaccurate interaction selection, and the attention contributions lost when tokens are discarded. Guided by this analysis, we propose Meta-Cached Sparse Attention (MC-Sparse), a training-free framework that selects individual key-value (KV) tokens while organizing similar queries into tile-aligned groups for efficient GPU execution. MC-Sparse caches metadata comprising query groups, KV indices selected using exact attention probabilities, and residuals between dense and sparse attention outputs, and reuses them across subsequent denoising steps. Across video and 3D generation models, MC-Sparse achieves higher fidelity to dense-attention outputs and larger denoising speedups than existing sparse-attention baselines, without visible quality degradation. Relative to dense attention, it delivers a $1.80\times$ denoising speedup on Minimax-H3-Base and a $2.32\times$ speedup on 3D asset generation, both with negligible quality loss.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.06801