Self-Evolving Coding Agents: From Digital Programs to Physical-World Intelligence
Researchers propose Physical Coding so robot agents store task state as revisable code and beat XR-1 on RoboCasa365.
Vision-language-action policies encode task state inside action sequences, so small layout or viewpoint changes cause failure. The authors propose Physical Coding: Code as World records objects, relations, constraints, and progress, while Code as Policy handles planning, verification, recovery, and execution. HexaAnything calls perception, planning, control, and VLA/WAM tools and stores verified traces as memory. On RoboCasa365 it improves Composite-Unseen and overall success versus XR-1, and a harness-trained HexaModel beats its base on every split; it also completes PhyBench experiments and most tasks on a dual-arm AgileX robot.
- Direct VLA policies fail when layouts, viewpoints, or instructions shift.
- Physical Coding splits Code as World state from Code as Policy.
- HexaAnything beats XR-1 VLA on RoboCasa365 Composite-Unseen and overall success.
- Harness-trained HexaModel wins every RoboCasa365 split over its base model.
- On PhyBench and an AgileX dual-arm robot, it finishes most tabletop tasks.
Full article253 words · extracted from huggingface.co · click to collapse
Vision-language-action (VLA) and world-action (WAM) models map observations and instructions directly to robot actions. This directness ties a policy to training: minor layout or viewpoint changes cause failure, and instructions generalize poorly. The root cause lies in representation: task requirements, conditions, progress, and failure recovery are implicitly encoded in action sequences, making them difficult to inspect or revise. Digital coding agents offer a precedent: LLMs call tools, verify results, and revise from feedback as executable code. The same working pattern of explicit state, manageable execution, and revisable procedures underlies generalization and long-horizon execution in the physical world, letting physical experience return as reusable programs, memory, or evidence. We propose Physical Coding, representing task state and execution as code. Code as World records objects, relations, constraints, and progress; Code as Policy organizes planning, verification, recovery, and execution. We build HexaAnything, which calls perception, planning, and control tools, including VLA/WAM policies, and makes in-the-loop decisions from external feedback. Verified traces become data and memory, enabling evolution from tools and Harness to model weights, architectures, and ultimately hardware and task design. On RoboCasa365, HexaAnything improves Composite-Unseen and overall success over XR-1 VLA, and its Harness-trained HexaModel beats the base on every split, indicating code traces internalize physical execution. On PhyBench and a dual-arm AgileX robot, the agent autonomously completes physics experiments and most tabletop tasks, often faster than published results. We observe data, model, and tool self-evolution; future work targets weight internalization, autonomous redesign of architectures, languages, representations, and tasks, and deployment in manufacturing and science.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.35432