Long-WAM: Scaling the Context of World-Action Models
Long-WAM raises robot task success by scaling causal world-action context with autoregressive video pretraining.
Long-WAM scales the context of causal world-action models for real-time robot control. Longer visual histories help mainly when the video foundation is pretrained autoregressively: on RoboCasa GR-1, context from 0.0 to 19.2 seconds raises success from 63.3% to 78.7%, while a bidirectional initialization shows no net gain. Long-WAM leads compared methods on LIBERO-Long, RoboTwin 2.0, and DOMINO, and reaches 95% success on dynamic cup stacking, where Pi0.5 and Fast-WAM succeed in none of 20 trials. On an RTX 5090, each action chunk, including future-video latent prediction, takes 107.4 ms, with deployment also shown on DGX Spark, Jetson AGX Thor, Unitree G1, and YAM.
- Autoregressive video pretraining makes longer history useful; bidirectional init does not.
- RoboCasa GR-1 success rises from 63.3% to 78.7% at 19.2 seconds of context.
- Dynamic cup stacking hits 95% success; Pi0.5 and Fast-WAM fail all 20 trials.
- Action chunks take 107.4 ms on RTX 5090, including future-video latent prediction.
Full article216 words · extracted from huggingface.co · click to collapse
Real-time robot control demands enough visual history to infer motion and task progress, but processing that history can delay action. We present Long-WAM, a model-system framework for scaling the context of causal world-action models under real-time control constraints. Our central finding is that access to history is not the same as using it: longer histories pay off far more when the video foundation is pretrained autoregressively (AR). We first learn causal prediction from robot and egocentric videos without action labels, then preserve this history-to-future structure during world-action adaptation. On RoboCasa GR-1, increasing context from 0.0 to 19.2 seconds raises success from 63.3% to 78.7%, whereas a bidirectionally pretrained initialization shows no net gain; robot-domain AR pretraining further raises peak success on GR-1 and LIBERO-Long. Long-WAM also achieves the best results among compared methods on LIBERO-Long, RoboTwin 2.0, and DOMINO. Streaming observation encoding, asynchronous execution, and hardware-specific acceleration enable deployment on RTX 5090, DGX Spark, and Jetson AGX Thor without dropping future prediction; on RTX 5090, each action chunk, including future-video latent prediction, takes 107.4 ms. Real-time deployment on Unitree G1 and YAM supports dynamic and long-horizon manipulation, including 95% success on dynamic cup stacking, where Pi0.5 and Fast-WAM succeed in none of 20 trials. As a memory-informed executor, Long-WAM also complements higher-level planning in composite tasks.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.10528