Reconstruct, Practice, Go Real: Guided Self-Improvement for Embodied Agents
RPG improves robot manipulation from 28.6% to 95% success in simulation and succeeds in all 30 physical trials.
Reconstruct, Practice, Go Real (RPG) improves robot execution without updating model weights by mining manipulation skills from an offline dataset, building related simulation practice tasks, and diagnosing failures with execution feedback, privileged simulator state, and dataset videos. It adds reusable symbolic skills, refines existing ones, and revises the system prompt, retaining only changes that pass cross-task checks. On held-out initializations of 22 manipulation tasks, success rose from 28.6% after the first practice round to 95.0% after 15 rounds, above ASPIRE at 75.5% and CaP-Agent0 with GPT-6 Astra Pro at 60.0%. After calibration, the frozen system succeeded in all 30 physical trials, ten each on three tasks.
- RPG improves robots via simulation practice without changing model weights.
- It writes reusable symbolic skills and revises the system prompt from failure diagnoses.
- Success on 22 held-out tasks rose from 28.6% to 95.0% over 15 rounds.
- It beat ASPIRE (75.5%) and CaP-Agent0 with GPT-6 Astra Pro (60.0%).
- After hardware adaptation, all 30 physical trials on three tasks succeeded.
Full article192 words · extracted from arxiv.org · click to collapse
Building reliable robot capabilities across diverse tasks requires substantial human effort to develop and maintain skills, design rewards, and integrate perception with control. We present Reconstruct, Practice, Go Real (RPG), a framework for autonomous improvement of robot execution systems without updating model weights. RPG identifies manipulation capabilities in an offline dataset and constructs related practice tasks in simulation. During practice, RPG uses execution feedback, privileged simulator state, and available dataset videos to diagnose failures. It develops new reusable symbolic skills, refines existing skills, and revises the system prompt based on these diagnoses. Cross-task evaluation tests individual candidate changes and merged revisions before they are retained for reuse. At test time, a multimodal LLM uses the resulting system prompt and skill library to coordinate perception and robot control. On held-out initializations of 22 manipulation tasks, RPG improves task success from 28.6% after the first practice round to 95.0% after 15 rounds, outperforming all evaluated baselines, including ASPIRE (75.5%) and CaP-Agent0 powered by GPT-6 Astra Pro (60.0%). After a common calibration and hardware-adaptation procedure, the frozen system succeeds in all 30 physical trials, with ten trials on each of three tasks. Project Website: https://rpg-robot.github.io/
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.02204