Adaptive Latent Capacity for World Models
ALeWM learns compact latent prefixes for JEPA world models and beats fixed-width LeWM.
Adaptive LeWorldModel (ALeWM) is a joint-embedding predictive architecture world model that concentrates predictive information in compact prefixes of a wide latent representation. It learns a sequence-conditioned distribution over prefix lengths and trains a predictor to estimate the next embedding from a sampled prefix. MixSIGReg regularizes masked embeddings toward a prior-weighted mixture with Gaussian active prefixes and zeros elsewhere, giving earlier coordinate blocks higher variance. In a controlled dynamical system and goal-conditioned visual control, ALeWM achieved higher mean success than tuned fixed-width LeWM while using lower planning capacity on average.
- ALeWM concentrates predictive information in compact latent prefixes.
- MixSIGReg uses Gaussian active prefixes and zeros later coordinates.
- Analysis places the most predictive information in earlier blocks.
- ALeWM beat fixed-width LeWM with lower average planning capacity.
Full article201 words · extracted from huggingface.co · click to collapse
We introduce Adaptive LeWorldModel (ALeWM), a world model based on a joint-embedding predictive architecture (JEPA) that learns to concentrate predictive information in compact prefixes of a wide latent representation. To encourage this ordering, ALeWM learns a sequence-conditioned distribution over prefix lengths and trains the predictor to estimate the full next embedding from a sampled input prefix. As standard anti-collapse objectives encourage variation across latent coordinates and do not organize them by predictive importance, we also introduce MixSIGReg. MixSIGReg regularizes the masked embeddings against a prior-weighted mixture with Gaussian active prefixes and zeros in the remaining coordinates. As a result, the ALeWM objective encourages early coordinates to retain information useful for prediction and recursive planning. Our analysis shows that the mixture distribution used by MixSIGReg assigns higher variance to earlier coordinate blocks and lower variance to later ones. In addition, we show that, under specified assumptions, prediction error is minimized by placing the information most useful for prediction in earlier blocks. Empirically, we study the behavior of ALeWM in a controlled dynamical system with known state variables and in goal-conditioned visual control. We show that ALeWM consistently achieves higher mean success rates than tuned fixed-width LeWM, with lower planning capacity on average.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.32921