Depth as Time in One-Step Generative Models
Depth-as-time analysis shows one-step diffusion models reorganize temporal denoising across network depth, enabling 16.6x compression of MeanFlow SiT-L/2.
The authors observe that the denoising computation multi-step diffusion performs across sampling steps unfolds across the depth of a single forward pass and can be recovered by decoding intermediate layers with the model's own output head. In MeanFlow, probing shorter transport intervals reveals both denoising and renoising within a single network evaluation, while drifting models without a time-indexed transport task lack this depthwise behavior. Treating layerwise computation explicitly as a flow, a single time-conditioned block can be trained to denoise across layers, compressing a MeanFlow SiT-L/2 model by 16.6x in parameters. The results suggest one-step generation reorganizes rather than eliminates diffusion's temporal computation.
- Denoising trajectory of diffusion unfolds across single-pass network depth
- MeanFlow shows denoise-then-renoise behavior within one evaluation
- Drifting models without time-indexed transport lack depthwise denoising
- Single time-conditioned block compresses MeanFlow SiT-L/2 by 16.6x
Full article263 words · extracted from arxiv.org · click to collapse
The recent wave of one-step generative models, which compress the multi-step trajectory of diffusion via either distillation or learned flow maps, has reached an inflection point where they can generate high-quality images. Here, we ask a natural question that follows from these advances: what happens to the denoising trajectory of multi-step diffusion when generation is compressed into a single forward pass? We offer an empirical observation we call \textit{depth as time}: the denoising computation that multi-step diffusion performs across sampling steps appears to unfold across the depth of a single forward pass, and can be recovered by decoding intermediate layers with the model's own output head. Most interestingly, we show that this depthwise computation depends on the transport task a flow map is trained to solve. The most surprising case is MeanFlow, where probing shorter transport intervals reveals both denoising and renoising within a single network evaluation. In contrast, generators trained without a time-indexed transport task, such as drifting models, do not exhibit the same depthwise denoising. Consequently, we show that models that exhibit the depthwise denoising phenomenon are more compressible across the layerwise computation: a MeanFlow \texttt{SiT-L/2} model can be compressed by $16.6\times$ in parameters into a single time-conditioned block. We offer an explanation for this denoise-then-renoise behavior and show that, when we treat the layerwise computation explicitly as a flow, a single time-conditioned block can be trained to denoise across layers, compressing a MeanFlow \texttt{SiT-L/2} model by $16.6\times$ in parameters. Together, these results suggest that the temporal computation of diffusion is not eliminated by one-step generation, but reorganized across network depth.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.03626