RealtimeWAM and Long-WAM advance robot world-action models
RealtimeWAM speeds inference about 25x with under 1% loss; Long-WAM then raises success by scaling autoregressive context.
RealtimeWAM distills multi-step world-action denoising into accurate one-step actions with Teacher-Anchored Consistency Distillation and overlaps video and action experts through Cross-Expert Wavefront Pipelining and block-wise KV-cache sharing. On LIBERO, LIBERO-Plus, and RoboTwin, Fast-WAM and Faster-WAM lose under 1% while reaching about 25x end-to-end speedup on an H100, and code and checkpoints were released. Long-WAM then scales causal world-action context for real-time control: on RoboCasa GR-1, history from 0 to 19.2 seconds lifts success from 63.3% to 78.7% only when the video foundation is pretrained autoregressively, not with bidirectional initialization. It leads compared methods on LIBERO-Long, RoboTwin 2.0, and DOMINO, and reaches 95% success on dynamic cup stacking, where Pi0.5 and Fast-WAM succeed in none of 20 trials. Action chunks, including future-video latent prediction, take 107.4 ms on an RTX 5090. Sources agree on those results but disagree on hardware scope: the Oct. 6 report lists deployment on DGX Spark, Jetson AGX Thor, Unitree G1, and YAM, while the Oct. 7 report names only Unitree G1 and YAM and associates the 95% cup-stacking result with that deployment.
- RealtimeWAM (reported 2026-10-04) uses Teacher-Anchored Consistency Distillation for one-step actions and Cross-Expert Wavefront Pipelining with block-wise KV-cache sharing.
- On LIBERO, LIBERO-Plus, and RoboTwin, Fast-WAM and Faster-WAM lose under 1% performance while reaching about 25x end-to-end speedup on an H100; code and checkpoints were released.
- Long-WAM (2026-10-06 and 2026-10-07) raises RoboCasa GR-1 success from 63.3% to 78.7% when context grows from 0 to 19.2 seconds under autoregressive video pretraining; bidirectional initialization shows no net gain.
- Long-WAM leads compared methods on LIBERO-Long, RoboTwin 2.0, and DOMINO.
- On dynamic cup stacking, Long-WAM reaches 95% success while Pi0.5 and Fast-WAM succeed in none of 20 trials.
- Each Long-WAM action chunk, including future-video latent prediction, takes 107.4 ms on an RTX 5090.
- Deployment lists differ: the Oct. 6 report names DGX Spark, Jetson AGX Thor, Unitree G1, and YAM; the Oct. 7 report names only Unitree G1 and YAM and ties the 95% result to that deployment.
Coverage timelineoldest first · each row is one article
- · 4d agoRealtimeWAM: One-Step Asynchronous World Action Models
Hugging Face daily papers· 54
RealtimeWAM enables one-step action generation and asynchronous inference, with about 25x speedup and under 1% drop.
- · 2d agoLong-WAM: Scaling the Context of World-Action Models
Hugging Face daily papers· 54
Long-WAM raises robot task success by scaling causal world-action context with autoregressive video pretraining.
- · 1d agoLong-WAM: Scaling the Context of World-Action Models
arXiv cs.AI / cs.LG / cs.CL· 48
Long-WAM scales robot world-action context, lifting RoboCasa success from 63.3% to 78.7%.