RealtimeWAM: One-Step Asynchronous World Action Models
RealtimeWAM enables one-step action generation and asynchronous inference, with about 25x speedup and under 1% drop.
World Action Models reuse video representations for action prediction, but multi-step denoising and sequential expert execution slow inference. RealtimeWAM uses Teacher-Anchored Consistency Distillation for accurate one-step actions and Cross-Expert Wavefront Pipelining to overlap the experts via block-wise KV-cache sharing. On LIBERO, LIBERO-Plus, and RoboTwin, variants including Fast-WAM and Faster-WAM lose under 1% performance while reaching about 25x end-to-end speedup on an H100. Code and checkpoints are released.
54