Beyond Future Prediction: Denoising as Generative Adaptation for Robot Control
NowWAM adapts diffusion transformers for robot control by denoising current frames, hitting 87.7% on LIBERO-Plus.
NowWAM adapts pretrained generative Diffusion Transformers for robot control by denoising the current observation and predicting actions from the same visual stream, without a future visual target. Matched experiments show past and future targets perform similarly, while training only at the clean endpoint hurts robustness. On LIBERO-Plus, NowWAM with FLUX2-Klein reaches 87.7%, 6.1 points above future-target co-training, while cutting visual tokens from 784 to 392 and step time from 2.85s to 1.63s. A text-to-image Z-Image backbone reaches 87.8%, indicating the gain is not limited to video or image-editing models.
- NowWAM denoises the current observation and predicts actions from the same visual stream.
- Past and future visual targets perform similarly under matched settings.
- Training only on the clean endpoint substantially reduces robustness.
- FLUX2-Klein reaches 87.7% on LIBERO-Plus, 6.1 points above the future-target baseline.
- Visual tokens drop from 784 to 392 and step time from 2.85s to 1.63s.
Full article215 words · extracted from huggingface.co · click to collapse
Pretrained generative Diffusion Transformers (DiTs) capture rich pixel-level visual and language-conditioned structure through large-scale image and video generation training. A growing line of robot policies builds on this generative prior, but how it should be transferred to control remains unclear, and existing approaches commonly instantiate this transfer through future visual prediction. We ask a more basic question: what a pretrained generative DiT actually contributes to action learning, and how this prior should be adapted for control. We introduce NowWAM, a future-target-free co-training formulation that denoises the current observation and predicts robot actions from the same visual stream, directly coupling the native generative objective to the action-facing representation across the denoising trajectory. Under matched controlled settings, past and future visual targets perform comparably, while restricting training to the clean endpoint substantially reduces robustness, suggesting that a separate future target is not essential for generative adaptation, while the denoising trajectory remains an effective interface for control. On LIBERO-Plus, NowWAM reaches 87.7% with FLUX2-Klein, improving over the future-target co-training baseline by 6.1 points while halving training visual tokens (784 to 392) and reducing step time from 2.85 s to 1.63 s, a 1.8x speedup. With the pure text-to-image Z-Image backbone, NowWAM further reaches 87.8%, showing that strong control adaptation is not tied to video generation or image-editing backbones.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.28339