Beyond Future Prediction: Denoising as Generative Adaptation for Robot Control
NowWAM adapts diffusion transformers for robot control by denoising current frames, hitting 87.7% on LIBERO-Plus.
NowWAM adapts pretrained generative Diffusion Transformers for robot control by denoising the current observation and predicting actions from the same visual stream, without a future visual target. Matched experiments show past and future targets perform similarly, while training only at the clean endpoint hurts robustness. On LIBERO-Plus, NowWAM with FLUX2-Klein reaches 87.7%, 6.1 points above future-target co-training, while cutting visual tokens from 784 to 392 and step time from 2.85s to 1.63s. A text-to-image Z-Image backbone reaches 87.8%, indicating the gain is not limited to video or image-editing models.