DeltaWAM: Delta World Action Models for Bimanual Manipulation
DeltaWAM predicts visual deltas and robot actions, raising RoboTwin success while cutting training and inference cost.
World-action models transfer video-generator priors to robot control but waste compute predicting largely unchanged future frames. DeltaWAM jointly predicts visual deltas and actions with dense-anchor, sparse-delta, and action streams, and Streaming Delta Memory updates cached anchors from compact observed deltas. On RoboTwin, DeltaWAM with SDM lifts average success over Fast-WAM from 81.3% to 85.4% clean and from 75.8% to 83.9% under visual randomization. The three architectures cut training FLOPs by 17.78-23.77%, while SDM cuts one-step inference latency 36.57% and FLOPs 31.55%.