SpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
SpaceCast-Bench tests predictive spatial reasoning; the best of 21 models scores 58.0% versus 87.2% for humans.
SpaceCast-Bench evaluates predictive spatial reasoning with 3,862 questions drawn from 182 real-world scenes across 16 task types and three levels: static perception, local prediction, and global prediction. Among 21 models, the strongest reaches 58.0% against 87.2% human performance, while spatially specialized models remain near chance. Bridge views help integrate distributed observations, and explicit 3D evidence helps more reliably than generated outcome images or videos. Fine-tuning on programmatically generated data raises Qwen3-VL-4B from 34.0% to 65.7%, with gains on six out-of-domain benchmarks.