Transferring the Intelligence of VLMs to Robotic Control
RoboDawn lets a VLM control robots via discrete commands, beating trained policies zero-shot and one-shot.
RoboDawn lets an agentic vision-language model control a robot through discrete translation, rotation, and gripper commands in a closed visual loop. A few in-context demonstrations ground both the interface and the task strategy, without task-specific robot training. On RoboTwin 2.0 C2R, success rises from 53.2% zero-shot to 73.6% one-shot, above the pi0.5 baseline at 46.0%. On RoboDojo it rises from 35.67% to 47.17%, and the same setup transfers to a Franka robot for block tasks.
52