Recursive Harness Distillation across Agents for Robot Manipulation
Recursive harness distillation lifts real-world robot manipulation success from 37.3% to 64.0%.
Recursive Harness Distillation accumulates robot-manipulation experience as a reusable playbook rather than updated model weights. A strong agent distills interventions for a light agent, then recursively revises the playbook using the light agent's execution feedback. In real-world manipulation, success rises from 37.3% to 64.0%. On SimplerEnv Bridge, the light agent with the playbook reaches 66.7% versus 41.7% for a GR00T-only baseline, and the strong agent using the same playbook reaches 79.2%.
- A strong agent distills interventions into a playbook for a lighter agent.
- The playbook is refined from the light agent's execution feedback.
- Real-world manipulation success rises from 37.3% to 64.0%.
- On SimplerEnv Bridge, the light agent reaches 66.7% versus 41.7% for GR00T-only.
- The same playbook lifts the strong agent to 79.2% without parameter updates.
Full article174 words · extracted from huggingface.co · click to collapse
A central goal in robotics is to enable manipulation across changing tasks and environments. Vision-language-action (VLA) models provide broad manipulation capabilities but can struggle when execution requires diagnosing failures and adapting behavior. Strong agents can discover effective interventions through interaction with these policies. We propose Recursive Harness Distillation to accumulate this experience as reusable guidance across agents. A strong agent distills its experience into a playbook for a light agent, then recursively refines the playbook using the light agent's execution feedback. The resulting playbook enables agents to reuse accumulated intervention knowledge in new task instances without updating model parameters. In real-world manipulation, the harness improves success from 37.3% to 64.0%. On SimplerEnv Bridge, the light agent with the playbook achieves 66.7% success, compared with 41.7% for the GR00T-only baseline, and outperforms the strong agent without a playbook. The same playbook also benefits the strong agent, which reaches 79.2% success. These results demonstrate the feasibility of harness distillation for robotics: intervention experience can be accumulated, refined through execution, and reused across agents to improve manipulation.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.33378