Long-WAM: Scaling the Context of World-Action Models
Long-WAM scales robot world-action context, lifting RoboCasa success from 63.3% to 78.7%.
Long-WAM scales the context of causal world-action models under real-time robot-control constraints. On RoboCasa GR-1, extending context from 0 to 19.2 seconds raises success from 63.3% to 78.7% when the video foundation is pretrained autoregressively, while a bidirectional initialization shows no net gain. It reports the best compared results on LIBERO-Long, RoboTwin 2.0, and DOMINO, with each action chunk taking 107.4 ms on an RTX 5090. Deployed on Unitree G1 and YAM, it reaches 95% success on dynamic cup stacking, where Pi0.5 and Fast-WAM succeed in none of 20 trials.