VISTA: A Visual Harness for Reasoning in an Interactive World
VISTA's visual-memory harness lifts Claude Opus 5.0 to a perfect ARC-AGI-3 score using 57.4% fewer actions.
VISTA is a visual harness that lets a general-purpose multimodal model observe an environment, keep past frames in lossless visual memory, and retrieve or reorganize that input while reasoning. On ARC-AGI-3 it raises Claude Opus 5.0's Relative Human Action Efficiency from 40.68 to 100.00, completing all 25 public games with 57.4% fewer actions than first-time human participants. With minimal adaptation, the same design substantially outperforms minimal harnesses on three additional visual game and puzzle benchmarks using the same underlying model.
- VISTA stores lossless visual memory and lets the model retrieve past observations.
- Claude Opus 5.0's ARC-AGI-3 efficiency score rises from 40.68 to 100.
- The model completes all 25 public games with 57.4% fewer actions than humans.
- The same harness beats minimal baselines on three further visual benchmarks.
Full article173 words · extracted from arxiv.org · click to collapse
We show that multimodal models possess strong reasoning abilities and that an appropriate harness can unlock their potential to solve tasks across diverse interactive environments. We introduce VISTA, a visual harness that gives a general-purpose multimodal model long-horizon vision. VISTA allows the model to directly perceive the environment through visual observations and maintains a lossless visual memory that preserves past observations in their original form. The model can actively retrieve these observations and reorganize its visual input as it reasons. On ARC-AGI-3, VISTA improves Claude Opus 5.0's Relative Human Action Efficiency score from 40.68 to a perfect 100.00, with the model completing all 25 public games using 57.4% fewer actions than first-time human participants. VISTA's simple design also allows it to extend naturally to diverse visual environments with minimal adaptation. Across three additional benchmarks covering a diverse range of visual games and puzzles, it substantially outperforms baselines using the same underlying model with minimal harnesses. Our results highlight VISTA's potential as a general-purpose visual harness for advancing multimodal agents in complex visual environments.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.02200