EmbodiedRSI: Active Continual Robot Learning Through Hypothesis-Guided Co-Evolution
EmbodiedRSI reaches 77% success on RoboCasa365, nearly doubling the best baseline.
EmbodiedRSI is a self-evolving agentic harness that chooses physical robot experiments to distinguish competing code and skill hypotheses, then turns outcomes into improved skills. A Fast-Slow architecture uses a Hypothesis Graph, Value-of-Information experiment selection, and reward-grounded hierarchical memory. On RoboCasa365 it reaches 77.0% overall success and 71.3% on Composite-Unseen, versus 40.1% for the best baseline, plus 86.8% overall on LIBERO-Pro. It also transfers zero-shot to a real robot at 71.3% overall success.
- Fast-Slow system maintains competing code and skill hypotheses in a graph.
- Value-of-Information selection picks physical trials that distinguish hypotheses.
- RoboCasa365 success is 77.0% versus 40.1% for the best baseline.
- Zero-shot real-robot transfer reaches 71.3% across challenging tasks.
Full article189 words · extracted from arxiv.org · click to collapse
Robot foundation models provide strong visuomotor control, yet their performance can degrade when object positions or task instructions change. Further improvements often require post-training on substantial robot data, which can be costly to collect through methods such as teleoperation. Agentic harnesses can adapt around the model, but current self-evolving harnesses use robot trials inefficiently when deciding which code and skill changes to pursue. We introduce EmbodiedRSI, a self-evolving agentic harness that autonomously decides where to explore next and turns the resulting physical interaction into improved code and skills. EmbodiedRSI realizes this through a Fast-Slow Dual-System Architecture, in which competing code and skill hypotheses are maintained in a Hypothesis Graph. Value-of-Information Experiment Selection chooses physical experiments that can distinguish these hypotheses. Their outcomes guide Code-Skill Co-Evolution. The Slow System builds Hierarchical Memory, and Reward-Grounded Memory Learning selects effective memory according to their value for later Fast-System improvement. On RoboCasa365, EmbodiedRSI reaches 77.0% overall success and 71.3% on Composite-Unseen, compared with 40.1% for the best baseline. EmbodiedRSI also reaches 86.8% overall success on LIBERO-Pro. Beyond benchmark performance, EmbodiedRSI transfers zero-shot to real-world robot, achieving 71.3% overall success across multiple challenging tasks.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.10498