SpaceCast-Bench Finds VLMs Weak on Spatial Prediction
SpaceCast-Bench finds the best of 21 vision-language models scores 58.0% versus 87.2% for humans.
SpaceCast-Bench evaluates predictive spatial reasoning in vision-language models using 3,862 questions drawn from 182 real-world scenes and 16 task types at static perception, local prediction, and global prediction levels. Among 21 models, the strongest reaches 58.0% compared with 87.2% human performance, while spatially specialized models stay near chance. Analyses find bridge views important for integrating observations and explicit 3D evidence more reliable than generated outcome images or videos. Fine-tuning on programmatically generated data lifts Qwen3-VL-4B from 34.0% to 65.7%, with macro-average gains on six out-of-domain benchmarks. The two reports agree on these figures and do not state conflicting results.
- SpaceCast-Bench tests predictive spatial reasoning with 3,862 questions from 182 real-world scenes across 16 task types.
- Tasks cover three levels: static perception, local prediction, and global prediction.
- Of 21 models, the strongest scores 58.0%, versus 87.2% for humans; spatially specialized models remain near chance.
- Bridge views help combine distributed observations, and explicit 3D evidence is more reliable than generated outcome images or videos.
- Fine-tuning on programmatic data raises Qwen3-VL-4B from 34.0% to 65.7%, with macro-average gains on six out-of-domain benchmarks.
Coverage timelineoldest first · each row is one article
- · 1d agoSpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
Hugging Face daily papers· 55
SpaceCast-Bench finds VLMs weak at predictive spatial reasoning; the best model scores 58 percent versus 87.2 percent for humans.
- · 15h agoSpaceCast-Bench: Evaluating Predictive Spatial Reasoning in Vision-Language Models
arXiv cs.AI / cs.LG / cs.CL· 56
SpaceCast-Bench tests predictive spatial reasoning; the best of 21 models scores 58.0% versus 87.2% for humans.