ZeroHour
Story · 1 source · 1 articlefirst updated ()

RoboSPA: New 527K-Trajectory Benchmark Shows VLA Models Struggle With Spatial Reasoning and Long-Horizon Planning

infoAI researchimportance 28
What's new: First merged summary for this story; no prior summary exists. The two reports cover the same RoboSPA paper and agree on all key figures (527K trajectories, 280 task variants, 56 base tasks, 10 task categories, five difficulty levels) and on the finding that current VLA models struggle with spatial relations and long-horizon planning. No substantive disagreements between sources were found.
Merged summary · glm-5.3-flash · rewritten as coverage arrives

RoboSPA introduces a large-scale robotic manipulation benchmark with 527K trajectories and 280 task variants, showing that current vision-language-action (VLA) models fail at complex spatial relations, precise low-level execution, and memory-intensive…

RoboSPA is a large-scale robotic manipulation dataset and benchmark for diagnosing embodied reasoning in vision-language-action models, covering fine-grained spatial reasoning and long-horizon procedural planning. The benchmark spans 10 task categories and 56 base tasks instantiated across five difficulty levels, yielding 280 task variants, with 527K trajectories collected across multiple embodiments and diverse scenes. It introduces diagnostic metrics that go beyond binary task success rates. Experiments on representative VLA models show that current systems struggle with complex spatial relations, precise low-level execution, and memory-intensive planning, including memory-heavy long-horizon planning. The two source reports (a Hugging Face daily papers entry dated 2026-09-03 and an arXiv cs.AI/cs.LG/cs.CL entry dated 2026-09-04) describe the same paper and agree on all substantive figures; no conflicts were found between them.

  • RoboSPA contains 527K trajectories collected across multiple embodiments and diverse scenes (both reports agree on the 527K figure).
  • The benchmark defines 280 task variants drawn from 56 base tasks instantiated across five difficulty levels.
  • It spans 10 task categories covering fine-grained spatial reasoning and long-horizon procedural planning.
  • RoboSPA introduces diagnostic metrics that go beyond binary task success rate.
  • Experiments show current VLA models struggle with complex spatial relations, precise low-level execution, and memory-intensive planning.
  • Report dates: Hugging Face daily papers entry dated 2026-09-03; arXiv (cs.AI/cs.LG/cs.CL) entry dated 2026-09-04.
ProductsRoboSPA
AI modelsRoboSPA

Coverage timeline

  1. · 12d ago
    Hugging Face daily papers· 28
    RoboSPA: Can VLA Models Go Beyond Simple Scenes and Short-Horizon Tasks?

    RoboSPA introduces a 527K-trajectory benchmark with 280 task variants showing current VLA models struggle with spatial reasoning and long-horizon planning.