LeWAM: A JEPA World Action Model with Diffusion-Steering-Based MPC
LeWAM is a decoder-free JEPA world action model that predicts dynamics and policies while planning in policy-head noise space.
LeWAM is a bidirectional transformer trained end-to-end for forward, backward, inverse-dynamics, and policy prediction on a decoder-free JEPA latent. Linear probes recover robot and object state better than a forward-only Le World Model, and the latent ignores visual distractors far better than a reconstruction-based world action model. Closed-loop control matches a same-size flow-matching policy, while model-predictive control in the policy head's noise space improves performance over sampling raw actions.
- Bidirectional transformer learns four prediction modes on a decoder-free JEPA latent.
- State probes beat a forward-only LeWM while ignoring visual distractors.
- Closed-loop control matches a same-size flow-matching policy.
- Noise-space MPC outperforms planning by sampling raw actions.
Full article153 words · extracted from arxiv.org · click to collapse
World action models (WAMs) predict actions and future observations, typically from a reconstruction-based representation that carries noisy, redundant information which can complicate downstream predictions. We introduce LeWAM, a bidirectional transformer for forward, backward, inverse dynamics and policy prediction, on a decoder-free JEPA latent trained end-to-end through all four modes. We see the following benefits: 1) Alignment: linear probes read robot and object state from LeWAM's latent better than from a regular Le World Model (a forward-only JEPA world model), while the latent ignores visual distractors as well as LeWM does and far better than a reconstruction-based WAM. 2) Acting: Closed-loop evaluations of LeWAM match a regular flow-matching policy trained on the same encoder at matched size, while also providing a world model. 3) Planning: Sampling raw actions when planning with WAMs lets MPC exploit dynamics-model inaccuracies; planning in the noise space of the policy head instead improves the closed-loop performance of these WAMs.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.12407