Papers report low-rank VLA updates and StructRL rewards
Two papers describe low-rank RL updates in VLA timestep modules and StructRL's verifiable subtask rewards.
Two Hugging Face daily papers from 2026-09-27 examine reinforcement learning for vision-language-action policies but report different findings. One finds that RL on flow-based models, including pi0.5 and GR00T N1.5/N1.6, produces substantially lower-rank updates concentrated in the action expert's Timestep Modules across LIBERO, ManiSkill, MetaWorld, and CALVIN. Module-replacement experiments show those modules account for a disproportionate share of RL gains and specialize to discrete denoising timesteps; shift-vector directions predict task success with ROC-AUC up to 99.6%, reflect cross-task transfer, and can steer policies to further gains without additional RL. Separately, StructRL replaces sparse terminal rewards with verifiable intermediate rewards that are granted only after ordered prerequisite subtasks are completed and scaled by completion pace. On RoboCasa365 and LIBERO-Long, with GR00T-N1.5 and pi 0.5, it consistently beats the online RL baselines tested, and Amazon Science released the code. The reports do not contradict each other; model names differ slightly in spelling, and the results are not presented as the same experiment.
- A low-rank-structure paper finds RL post-training of flow-based vision-language-action models produces substantially lower-rank updates concentrated in the action expert's Timestep Modules.
- That study evaluates pi0.5 and GR00T N1.5/N1.6 on LIBERO, ManiSkill, MetaWorld, and CALVIN; module-replacement tests attribute a disproportionate share of RL gains to those modules.
- Shift-vector update directions predict task success with ROC-AUC up to 99.6%, reflect cross-task transfer, and steering along them improves policies without further RL.
- StructRL is an online method that replaces sparse terminal rewards with intermediate rewards granted only after ordered prerequisite subtasks are verified complete, scaled by completion pace.
- On RoboCasa365 and LIBERO-Long, using GR00T-N1.5 and pi 0.5, StructRL consistently outperforms the online RL baselines evaluated.
- Amazon Science released StructRL code; both items were listed on Hugging Face daily papers on 2026-09-27.
Coverage timelineoldest first · each row is one article
- · 11d agoStructRL: Online Structured Reinforcement Learning for Long-Horizon Vision-Language-Action Tasks
Hugging Face daily papers· 45
StructRL uses verifiable intermediate subtask rewards to improve long-horizon vision-language-action policies.
- · 11d agoThe Low-Rank Structure of VLA Reinforcement Learning
Hugging Face daily papers· 46
RL post-training of flow-based VLA models produces low-rank updates concentrated in action-expert Timestep Modules.