The Low-Rank Structure of VLA Reinforcement Learning
RL post-training of flow-based VLA models produces low-rank updates concentrated in action-expert Timestep Modules.
Researchers find reinforcement learning on flow-based vision-language-action models, including pi0.5 and GR00T N1.5/N1.6, produces substantially lower-rank updates concentrated in the action expert's Timestep Modules. Tests span LIBERO, ManiSkill, MetaWorld, and CALVIN, and module-replacement experiments show these modules account for a disproportionate share of RL gains. RL specializes the modules to discrete denoising timesteps; shift-vector update directions predict task success with ROC-AUC up to 99.6% and reflect cross-task transfer. Steering along those directions further improves policies without additional RL.