FlexiWorld: Learning and Planning via Flexible Action Chunks Across Multiple Time Scales
FlexiWorld uses variable-length action chunks and mixed-span goals, reaching 89.29 percent mean success.
FlexiWorld is a JEPA-based world model trained on varying goal spans and randomly partitioned variable-length action chunks, with a causal action encoder and an autoregressive actor. Student Forcing trains on generated action prefixes to reduce exposure bias, and planning uses Actor-Residual Cross-Entropy Method (ARCEM). Across four benchmarks and goal distances, FlexiWorld with ARCEM reaches 89.29 percent mean success versus 83.98 percent for the strongest baseline. Longer planning chunks speed ARCEM by about 1.3x on average without retraining while keeping comparable success.
- JEPA world model uses mixed-span goals and variable-length chunks.
- ARCEM plans with residual search and autoregressive feedback.
- Mean success is 89.29 percent versus 83.98 percent for the best baseline.
- Longer chunks speed planning about 1.3x without retraining.
Full article179 words · extracted from arxiv.org · click to collapse
Latent world models predict future states for goal-directed planning using action chunks spanning multiple primitive steps. Existing methods typically use fixed-length chunks and either omit goal-conditioned action generation or limit their supervision to short goal spans. We introduce FlexiWorld, a JEPA-based world model that combines mixed-span goal supervision with variable-length action chunks to improve long-horizon control. During training, we sample varying goal spans and randomly partition the actions into variable-length chunks. We jointly train the world model with a causal action encoder that embeds variable-length chunks and an autoregressive actor that generates primitive actions sequentially. Student Forcing reduces exposure bias by training on generated action prefixes. For planning, Actor-Residual Cross-Entropy Method (ARCEM) combines action-residual search with within-chunk autoregressive feedback and chunk-boundary latent prediction. Across four benchmarks and goal distances, FlexiWorld with ARCEM achieves 89.29% mean success, compared with 83.98% for the strongest baseline. PushT ablations show improved direct control from mixed-span supervision, variable-length chunks, and Student Forcing. Without retraining, FlexiWorld supports different planning chunk lengths: longer chunks accelerate ARCEM by approximately $1.3\times$ on average while maintaining comparable average success.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.35138