Coding Agents for Generalized Task and Motion Planning Problems
Coding agents beat hand-engineered planners on generalized task and motion planning in simulation.
The paper tests whether coding agents can synthesize programs that generalize across task-and-motion-planning instances, replacing TAMP-specific engineering. Claude Code (Opus 5) and Codex (GPT-5.6 Sol and GPT-6 Astra) were given a task description and simulator access, then frozen programs were evaluated on unseen KinDER and PDDLStream environments with higher object counts. Across methods, 980 programs were each run on 100 held-out instances (98,000 episodes). All three agent setups beat hand-engineered planners, one-shot generation, and an LLM generalized-planning baseline, with mean success of 56% to 95% versus 47% where a planner exists, and they use about an order of magnitude less computation per instance as scenes grow.
- Agents write programs for generalized TAMP inside a fixed synthesis budget.
- Claude Code with Opus 5 and Codex with two GPT variants were tested.
- 980 programs ran on 100 held-out instances each, 98,000 episodes total.
- Mean success was 56% to 95%, versus 47% for available hand-engineered planners.
- Generated programs stay more successful as object counts grow and use less compute.
Full article265 words · extracted from arxiv.org · click to collapse
Task and motion planning (TAMP) problems remain difficult even with full observability and object-centric states because discrete decisions are tightly coupled to geometric, kinematic, and dynamic constraints. Generalized TAMP addresses this difficulty by exploiting regularities across problem instances to reduce planning effort on new instances. However, existing methods require substantial TAMP-specific engineering. We investigate whether coding agents can automate this process by synthesizing programs that generalize across instances. Given a task description and simulator access, each agent chooses how to interact with the environment while developing a program within a fixed synthesis budget. The program is then frozen and evaluated on unseen instances. We evaluate Claude Code (Opus 5) and Codex (GPT-5.6 Sol and GPT-6 Astra) on 28 simulated environments from KinDER and PDDLStream, with object counts beyond those evaluated in the original benchmark. Across all program synthesis methods, we evaluate 980 generated programs on 100 held-out instances each, 98,000 evaluation episodes in total. Overall, we find that coding agents are surprisingly effective at generalized TAMP: all three agent configurations outperform hand-engineered planners, one-shot generation, and an LLM-based generalized planning baseline in mean success (56% to 95% versus 47% for the planners, on the 16 environments where a planner is available). As object counts grow, the agents' programs maintain higher success than the planner, using an order of magnitude less computation per instance on average. Logs show agents using interaction to calibrate physical models, test edge cases, and refine strategies. We release all code, including the full prompts given to the agents. These findings suggest that coding agents are a strong baseline for generalized TAMP.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.30233