Grounded Action Model: 3D Grounding as a Foundation for Robotics
Grounded Action Models use 3D object grounding and beat prior robot policies on RoboTwin and LIBERO-PRO.
Grounded Action Models build robot policies around explicit 3D object grounding rather than learning metric location only from demonstrations. Language, point, or box prompts become a shared object-centric representation that a multi-stream transformer mixes with robot state to predict action chunks. On RoboTwin 2.0, GAM averages 55.3% success across 50 tasks and 47.6% under scene randomization. It reaches 61% on LIBERO-PRO and, on real robots, 17 of 20 successes under visual shift on a bimanual YAM, with a Molmo2 planner scoring 64.7% in-distribution step completion on a Franka.