SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
SpaceCast-Bench finds VLMs weak at predictive spatial reasoning; the best model scores 58 percent versus 87.2 percent for humans.
SpaceCast-Bench tests predictive spatial reasoning with 3,862 questions drawn from 182 real-world scenes, covering 16 task types at static perception, local prediction, and global prediction levels. Of 21 evaluated models, the strongest reaches 58.0% against 87.2% human performance, while spatially specialized models remain near chance. Analyses find bridge views important for combining observations and explicit 3D evidence more reliable than generated outcome images or videos. Fine-tuning on programmatic data raises Qwen3-VL-4B from 34.0% to 65.7%, with macro-average gains on six out-of-domain benchmarks.
- 3,862 questions from 182 scenes across 16 task types
- Best of 21 models scores 58.0% versus 87.2% for humans
- Explicit 3D evidence helps more than generated outcome images or video
- Fine-tuning lifts Qwen3-VL-4B from 34.0% to 65.7%
Full article165 words · extracted from huggingface.co · click to collapse
Existing spatial reasoning benchmarks mainly test spatial perception: reading off relations already visible in the input. Yet real-world spatial intelligence demands predictive spatial reasoning: constructing a scene from observations, anticipating how an intervention changes it, and reasoning about the unseen outcome. We introduce SpaceCast-Bench, the first benchmark to directly and diagnostically evaluate this capability. Built around an observe-transform-infer framework, its 3,862 questions from 182 real-world scenes span 16 task types at three levels: static perception, local prediction, and global prediction, progressively requiring scene understanding, spatial state updating, and relational inference over unobserved outcomes. Evaluating 21 models exposes a stark gap: the strongest model reaches only 58.0% against 87.2% human performance, while spatially specialized models remain near random chance. Controlled analyses further reveal that bridge views are critical for integrating distributed observations, and that explicit 3D evidence benefits models more reliably than generated outcome images or videos. Fine-tuning on our programmatically generated data lifts Qwen3-VL-4B from 34.0% to 65.7% with macro-average gains across six out-of-domain benchmarks.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.12402