RoboRSI: Stable, efficient, and reusable robot self-evolution in complex real-world environments
RoboRSI lets a mobile manipulator reuse validated skills, beating baselines by up to 11 points.
RoboRSI is a robot self-improvement system that organizes execution experience with Top-Down Skill Refinement. Tasks are split into compound, atomic, and base skills with explicit contracts, and revisions are confined to the branch responsible for a failure. Manager, Planner, Engineer, and Reviewer roles plan, execute, diagnose, and release validated skills. On a mobile manipulator it developed household cleanup over 104 rounds and beat the strongest baseline by 2.7 to 11.0 points on LIBERO, LIBERO-PRO, LIBERO-Plus, and RoboTwin.
- Top-Down Skill Refinement attributes each repair to the responsible skill branch.
- A mobile manipulator learned multi-object household cleanup over 104 rounds.
- Simulation success exceeds the strongest baseline by 2.7 to 11.0 points.
- Validated skill sequences are consolidated into reusable compound skills.
Full article197 words · extracted from arxiv.org · click to collapse
A generalist robot should not only perform diverse tasks but also improve through experience, turning what it learns during execution into capabilities that later tasks can reuse. Robot agents that act through code can already repair programs from execution feedback, yet it remains a central challenge to organize this experience around the task structure that gives it meaning, so that each repair is attributed to the responsible capability, supported by execution evidence, and validated before it is reused. We introduce RoboRSI, a robot self-improvement system built on Top-Down Skill Refinement (TSR). TSR decomposes tasks into compound, atomic, and base skills with scoped responsibilities and explicit input--output contracts, attributes each execution outcome to the responsible branch, and confines revision to that branch. Building upon this structure, a Manager, Planner, Engineer, and Reviewer coordinate planning, execution, diagnosis, and the validated release of new skills, while people steer the process through objectives and corrections; stable skill sequences are further consolidated into reusable compound skills. On a mobile manipulator, RoboRSI develops multi-object household cleanup over 104 rounds. In simulation, it achieves the highest success rate on LIBERO, LIBERO-PRO, LIBERO-Plus, and RoboTwin, exceeding the strongest baseline by 2.7 to 11.0 percentage points.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.12424