Rolling-WAM: World Action Models with Rolling Imagination
Rolling-WAM speeds robot replanning 4.5x by staggering joint video-action denoising across cycles.
World Action Models jointly denoise future video and actions for robotic manipulation, but full denoising at every replan adds latency. Rolling-WAM keeps a sliding window of video-action chunks at staggered noise levels, fully denoising only the imminent action chunk while partially refining farther chunks. Evaluations on LIBERO, RoboTwin, and a real Unitree G1 humanoid report competitive manipulation performance and a 4.5x steady-state replanning speedup over standard joint WAMs.
- Distributes joint video-action denoising across successive replanning cycles
- Sliding window of chunks at staggered noise levels
- 4.5x steady-state replanning speedup versus standard joint WAMs
- Competitive results on LIBERO, RoboTwin, and a Unitree G1 humanoid
Full article153 words · extracted from huggingface.co · click to collapse
World Action Models (WAMs) couple action generation with future visual prediction for robotic manipulation. However, completing the joint video-action denoising process at each replanning cycle incurs substantial latency, delaying action updates and limiting closed-loop responsiveness. We present Rolling-WAM, a formulation that distributes joint denoising across successive replanning cycles. Our method maintains a sliding window of video-action chunks at staggered noise levels. At each step, a rolling noise schedule fully denoises the imminent action chunk for execution, while partially refining farther-future chunks. As the window advances with new camera observations, the retained future chunks continue their denoising process. This distributes the computational cost over time while carrying an evolving visual-action context across chunk boundaries. Evaluations on LIBERO, RoboTwin, and a real-world Unitree G1 humanoid show that Rolling-WAM achieves competitive manipulation performance. By removing the need to denoise the entire prediction horizon from scratch, it delivers a 4.5x steady-state replanning speedup over standard joint WAMs.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.30247