In-Context Robot Learning with VLM Agentsnew
GPT-Policy uses commercial VLM agents like GPT-6 Astra for in-context robot learning, translating demonstrations and feedback into verified robot actions without gradient updates.
The paper (arXiv 2609.19138) presents GPT-Policy, a framework combining a context compiler that preserves task-relevant visual transitions, a VLM such as GPT-6 Astra that proposes robot-tool actions, and a constrained controller that verifies and executes each action. Real-robot trials show human video demonstrations improve task completion even without robot action labels, while aligned action references yield further gains on contact-sensitive tasks. Evaluation covers success and efficiency metrics, matched model comparisons, and controlled context ablations.