GameHorizon Suite: Multi-Horizon Data and Evaluation in Gameplay
GameHorizon provides 5,000 hours of AAA gameplay data and a benchmark covering 47 models.
GameHorizon Suite measures AI gameplay ability across multiple time horizons. Its dataset contains 5,000 hours of recordings from 21 AAA games, collected by 100 expert players, with aligned video, actions, and multi-horizon instructions. The benchmark supports reproducible offline tests and stepwise online evaluation. The authors tested 47 models through more than one million invocations and plan to release the dataset, annotator, and benchmark.
- Dataset has 5,000 hours from 21 games and 100 players
- Benchmark combines offline questions with stepwise online tests
- Forty-seven models received more than one million evaluations
- Authors plan to release data, annotator, and benchmark
Full article238 words · extracted from arxiv.org · click to collapse
Modern video games provide a measurable testbed for AI models, combining abilities of visual understanding, instruction decomposition, goal planning, and precise action control over multiple temporal horizons. Existing datasets and benchmarks, however, either cover a narrow range of games, lack language instructions, or rely on high-variance online rollouts. To address these challenges, we introduce GameHorizon, a unified data and evaluation suite that measures gameplay capabilities at different horizons for diverse model families. GameHorizon Suite consists of three components. First, GameHorizon-Annotator is a scalable and automated annotation pipeline for multi-horizon instructions. Second, utilizing the pipeline, we construct GameHorizon-Data, the first large-scale AAA gameplay dataset with temporally aligned videos, player actions, and multi-horizon instructions. It comprises 5,000 hours of recordings from 21 games, collected by 100 human expert players. Third, we build GameHorizon-Bench with reproducible offline and stepwise online testing. The offline track enables reproducible evaluation using thousands of standardized questions organized into three primary tasks and a series of diagnostic variants, while the online track tests whether offline scores reflect actual gameplay capabilities and localizes failures to specific steps within long-horizon gameplay. Based on our GameHorizon Suite, we evaluate 47 models through more than one million model invocations, revealing a meaningful hierarchy of task difficulty and pronounced differences in model capabilities. Our work can provide a standardized yardstick for evaluating gameplay capabilities across horizons and model families. We will release our dataset, annotator, and benchmark to facilitate future research.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.25001