RobotWorld: Benchmarking Multimodal Agents for Robot Use Across Diverse Tasks and Embodiments
RobotWorld shows multimodal agents still fail to reliably complete diverse robot tasks.
RobotWorld is a simulation testbed that turns instructions and observations into physical robot actions across 84 tasks spanning manipulation, mobile manipulation, locomotion, driving, and aerial control. Agents can assemble sophisticated perception and control workflows, including segmentation, calibration, spatial estimation, and dynamics computation, but these skills do not reliably compose into successful behavior. Agents lose object state, fail to correct ineffective actions, recover too late, or treat unfinished tasks as complete. Astra performs better on spatial and constrained-contact goals, while Opus 5.5 does better on continuous-balance and timed-interaction goals.
- RobotWorld has 84 simulated tasks across manipulation, locomotion, driving, and aerial control.
- Agents can build perception and control workflows but often fail to compose them.
- Common failures include lost object state, late recovery, and false task completion.
- Astra is stronger on spatial and contact goals; Opus 5.5 on balance and timed interaction.
Full article208 words · extracted from huggingface.co · click to collapse
General-purpose agents increasingly write code, use tools, and complete complex digital tasks, raising the question of how far these capabilities carry into the physical world. To investigate this, we introduce RobotWorld, a challenging simulation testbed for robot use: turning instructions and observations into physical task execution through robot interfaces. Its 84 tasks span manipulation, mobile manipulation, locomotion, driving, and aerial control, with explicit interaction budgets and executable success checks. By analysing task outcomes alongside execution traces, we identify both the capabilities that transfer and the gaps that prevent reliable completion. Furthermore, we find that current agents can construct sophisticated perception and control workflows, including image segmentation, camera calibration, spatial estimation, and dynamics-based computation. These capabilities, however, do not consistently compose into successful behaviour: agents lose task-relevant object states despite reaching commanded poses, fail to correct ineffective actions, recover too late, or mistake unfinished tasks for completion. This uneven transfer also differs across models: Astra succeeds more often on spatial and constrained-contact goals, whereas Opus 5.5 succeeds more often on continuous-balance and timed-interaction goals. By linking these outcomes to execution behaviour, RobotWorld provides both a rigorous proving ground and an empirical account of the remaining capability gaps, thereby establishing concrete targets for training and designing more reliable physical-world agents.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.10409