Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens
PyRUA-Lean raises GPT-6 Astra robot-agent success from 63.1% to 71.7% while cutting input tokens 65%.
Vision-language robot agents waste tokens on repeated calls and redundant observations. PyRUA-Lean lets a GPT-6 Astra planner compose classical primitives and vision-language-action policies into Python cells that retry locally and return only requested images and state. On 700 simulated tasks from LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, success rose from 63.1% to 71.7% under equal LLM-call budgets. On jointly solved instances it used 49% fewer LLM calls and 65% fewer input tokens.
- Success rose from 63.1% to 71.7% across 700 simulated tasks.
- Compared with a tool-calling baseline under equal LLM-call budgets.
- Jointly solved tasks used 49% fewer calls and 65% fewer tokens.
- Python cells compose primitives and policies, returning only requested feedback.
Full article128 words · extracted from huggingface.co · click to collapse
Vision language model (VLM) agents can control robots through visual feedback and action primitives, but repeated model invocations and redundant observations incur substantial token overhead. We introduce PyRUA-Lean, an interactive code-execution framework that couples feedback-driven primitive composition with selective observation: the agent composes classical robot primitives and learned vision-language-action (VLA) policies into Python cells that perform conditional checks and local retries, returning only explicitly requested images and state feedback for replanning. Across 700 simulated task instances from LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, we compare PyRUA-Lean with a tool-calling baseline using the same GPT-6 Astra planner and underlying robot primitives. Under equal LLM-call budgets, PyRUA-Lean increases overall success from 63.1% to 71.7%. On instances solved by both agents, it uses 49% fewer LLM calls and 65% fewer input tokens.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.01939