Two papers report separate robot self-evolution benchmark gains
EmbodiedRSI reports 77% RoboCasa365 success, nearly double a baseline, while RoboRSI says skill refinement beat baselines by up to 11 points.
Two arXiv reports from 7 and 8 October 2026 describe separate robot self-improvement systems, not one shared result. EmbodiedRSI keeps competing code and skill hypotheses in a graph and chooses physical trials by value of information; it reports 77.0% overall success on RoboCasa365 versus 40.1% for the best baseline, 71.3% on Composite-Unseen, 86.8% overall on LIBERO-Pro, and 71.3% zero-shot on a real robot. RoboRSI splits tasks into compound, atomic, and base skills and uses Manager, Planner, Engineer, and Reviewer roles so revisions stay on the branch responsible for a failure. On a mobile manipulator it developed household cleanup over 104 rounds and exceeded the strongest baseline by 2.7 to 11.0 points on LIBERO, LIBERO-PRO, LIBERO-Plus, and RoboTwin. The sources do not contradict a common metric: they name different methods and suites, and their LIBERO-Pro figures are not directly comparable.
- EmbodiedRSI (arXiv, 2026-10-07) is a Fast-Slow harness with a Hypothesis Graph, Value-of-Information experiment selection, and reward-grounded hierarchical memory.
- EmbodiedRSI reports 77.0% overall success on RoboCasa365 and 71.3% on Composite-Unseen, versus 40.1% for the best baseline.
- EmbodiedRSI reports 86.8% overall on LIBERO-Pro and 71.3% zero-shot success on a real robot.
- RoboRSI (arXiv, 2026-10-08) uses Top-Down Skill Refinement with compound, atomic, and base skills under explicit contracts.
- RoboRSI assigns Manager, Planner, Engineer, and Reviewer roles and confines each repair to the skill branch tied to a failure.
- A RoboRSI mobile manipulator developed multi-object household cleanup over 104 rounds and consolidated validated sequences into reusable compound skills.
- RoboRSI beat the strongest baseline by 2.7 to 11.0 points on LIBERO, LIBERO-PRO, LIBERO-Plus, and RoboTwin.
- The reports name different systems and do not give matching absolute scores; LIBERO-Pro results are an absolute 86.8% in one and only a margin over a baseline in the other.
Coverage timelineoldest first · each row is one article
- · 3d agoEmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution
arXiv cs.AI / cs.LG / cs.CL· 48
EmbodiedRSI reaches 77% success on RoboCasa365, nearly doubling the best baseline.
- · 2d agoRoboRSI: Stable, efficient, and reusable robot self-evolution in complex real-world environments
arXiv cs.AI / cs.LG / cs.CL· 38
RoboRSI lets a mobile manipulator reuse validated skills, beating baselines by up to 11 points.