Recursive Video In-Context Learning for Agentic Robot
RV-ICL lets robot agents navigate one demo video, raising LIBERO success without extra training.
Recursive Video In-Context Learning turns one demonstration into a hierarchy of sub-events such as grasps and releases, exposed through read-only tools rather than stuffed into the prompt. Levels run from whole-task keyframes down to phases, moments, and short clips. Built on RPent and training-free, the agent plans from coarse views and re-enters only the clip for the current sub-goal. Success rises from 92.6% to 96.5% on LIBERO-PRO and from 86.7% to 95.8% on LIBERO-Plus.
- One demonstration per task is enough and no extra training is required.
- The agent plans from coarse levels and loads only the current clip.
- LIBERO-PRO success rises from 92.6% to 96.5%.
- LIBERO-Plus success rises from 86.7% to 95.8%.
Full article192 words · extracted from arxiv.org · click to collapse
LLM agents that orchestrate frozen vision-language-action (VLA) policies improve across episodes through text memory, which records what the agent did but not how the task is done. A demonstration video shows it, but fits poorly into an agent's context. The full video slows every turn, fixed keyframes lose the contact detail that decides whether a grasp holds, and what the agent needs shifts from the task's structure while planning to the frames around each contact. We introduce Recursive Video In-Context Learning (RV-ICL), a training-free method that turns a demonstration into a hierarchy the agent navigates rather than a prompt it receives. The hierarchy is built from the sub-events of the demonstration, such as grasps and releases. Its levels grow finer, from keyframes of the whole task to phases, moments and short clips, and are exposed through read-only tools. The agent reads the coarse levels before planning. During execution it re-enters the hierarchy whenever a step needs more detail and loads only the clip of its current sub-goal. One demonstration per task is enough. Built on RPent, RV-ICL raises success from 92.6% to 96.5% on LIBERO-PRO and from 86.7% to 95.8% on LIBERO-Plus.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.06843