StructRL: Online Structured Reinforcement Learning for Long-Horizon Vision-Language-Action Tasks
StructRL uses verifiable intermediate subtask rewards to improve long-horizon vision-language-action policies.
StructRL is an online reinforcement-learning method for long-horizon vision-language-action tasks that replaces sparse terminal rewards with verifiable subtask rewards. Each task is decomposed into ordered subtasks; intermediate reward is granted only after prerequisites are completed and is scaled by completion pace. On RoboCasa365 and LIBERO-Long, using GR00T-N1.5 and pi 0.5, it consistently outperforms the online RL baselines evaluated. Code is released by Amazon Science.
- StructRL rewards only after prerequisite subtasks are verified complete.
- Rewards scale with how quickly each subtask is finished.
- Tests use RoboCasa365 and LIBERO-Long.
- Policies evaluated include GR00T-N1.5 and pi 0.5.
- It beats compared online RL baselines; code is on GitHub.
Full article141 words · extracted from huggingface.co · click to collapse
Vision-language-action (VLA) models perform well on shorter-horizon manipulation tasks but still struggle with long-horizon tasks that require multiple dependent manipulations from a single command. Online reinforcement learning (RL) can improve these policies through environment interaction, yet many existing methods provide reward only after the complete task succeeds. However, such terminal supervision is sparse and does not distinguish early failures from rollouts that make substantial partial progress. We propose StructRL, an online RL framework that constructs structured intermediate supervision from verifiable subtask completions. StructRL decomposes each task into verifiable subtasks, grants intermediate rewards only after the prerequisite subtasks have been completed, and scales each reward according to completion pace. Across RoboCasa365 and LIBERO-Long with GR00T-N1.5 and pi 0.5, StructRL consistently outperforms evaluated online RL baselines. These results show that verifiable, structured intermediate rewards improve long-horizon VLA post-training. Code is available at https://github.com/amazon-science/StructRL.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.36352