Game Arena: Strategic LLM Evaluation in Competitive Environments
Kaggle Game Arena ranks LLMs through ongoing Chess, Poker, and Werewolf competitions instead of static benchmarks.
The report introduces Kaggle Game Arena, an open platform for evaluating LLMs through competitive games rather than static benchmarks. Its three pilot environments are Chess, Poker, and Werewolf, covering perfect-information, imperfect-information, and multiplayer settings. For each game it describes the environment, metrics, and results from full model competitions. The infrastructure is designed for reproducible, ground-truth evaluation and for adding new games over time.
- Open Kaggle platform for head-to-head LLM matchups
- Pilot games are Chess, Poker, and Werewolf
- Covers perfect information, hidden information, and multiplayer play
- Competition strength is intended to rise as models improve
Full article133 words · extracted from huggingface.co · click to collapse
We introduce Kaggle Game Arena, an open and ever-expanding platform to evaluate large language models (LLMs) through competitive games. Different from static benchmarks, game arena enables models to play head-to-head matchups in structured environments where the gameplay strength naturally increases as models evolve, preventing performance saturation. This technical report details the infrastructure behind Game Arena and describes the three pilot game environments: Chess, Poker, and Werewolf. These environments span perfect information, imperfect information, and multiplayer game settings, enabling a systematic study of models' strategic planning, adaptation, and robustness under uncertainty. For each game, we provide a detailed description of the environment, evaluation metrics, and results from running full competitions across models. Through robust infrastructure and large-scale ground-truth based evaluation, Game Arena ensures reproducibility, transparency and generalizability to new games and variants over time.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.31473