RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement
RSIGame recursively improves generated games, letting Qwen3.8-27B beat GPT-5.5 one-shot on GameCraft-Bench.
RSIGame is an autonomous game-development framework that combines a local explore-diagnose-improve loop with a global loop that tracks quality, keeps the best checkpoint, and detects saturation or regression. Successful experience is also internalized through training rather than only test-time revision. Across 140 GameCraft-Bench tasks, two engines, and five generators, it improves quality under matched budgets. Trained Qwen3.8-27B reaches 61.38 on Godot and 58.53 on Phaser, exceeding GPT-5.5 one-shot scores while cutting Qwen generation tokens by 11 times.
- Local and global loops explore, diagnose, revise, and guard against regression.
- Evaluation covers 140 GameCraft-Bench tasks, two engines, and five generators.
- Trained Qwen3.8-27B scores 61.38 on Godot and 58.53 on Phaser.
- Those results beat GPT-5.5 one-shot while using 11 times fewer tokens.
Full article179 words · extracted from huggingface.co · click to collapse
Recent advances in large language models have made automatic game generation increasingly feasible, yet reliably improving generated games beyond a playable version remains challenging. Naive iterative refinement can easily overfit a small set of test cases, producing fragile games with unresolved bugs, missing behaviors, and poor generalization to broader player interactions. We introduce RSIGame, an autonomous agentic game development framework with recursive self-improvement. RSIGame organizes development into complementary local and global loops. Concretely, a local explore-diagnose-improve loop broadly explores the executable game, diagnoses and prioritizes discovered issues, and performs evidence-grounded revision, where an evolving checklist continually accumulates new testing and improvement guidance. A global loop tracks overall quality, preserves the best checkpoint, and detects saturation or regression over long-horizon development. Beyond test-time improvement, RSIGame further internalizes successful development experience into the generator through training. Across 140 GameCraft-Bench tasks, two game engines, and five generators, RSIGame consistently improves game quality under matched development budgets. Notably, experience internalization enables Qwen3.8-27B to reach 61.38 on Godot and 58.53 on Phaser, exceeding GPT-5.5 one-shot scores while reducing Qwen's generation tokens by 11 times.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.39045