Language Models that Play Chess and Explain Their Moves
Queen, a 4B-parameter chess-language model, reaches 2697 Elo via iterative distillation while explaining moves, beating frontier models with ~1000x fewer parameters.
Queen combines a silent expert chess encoder with an instruction-tuned language model through cross-attention in an encoder-decoder architecture, trained with a question-answering curriculum. An iterative distillation loop using a natural-language analog of the Bellman update adds over 900 Elo (1782 to 2697) across seven iterations, surpassing frontier models on playing strength and puzzle accuracy. The authors propose the recipe generalizes to domains with silent expert encoders such as games, robotics, and computer use.
- 4B chess-LM gains 900+ Elo to 2697 over seven distillation iterations
- Cross-attention integrates expert chess encoder with instruction-tuned LM
- Outperforms frontier models with three orders of magnitude fewer parameters
- Architecture claimed to generalize to robotics and computer-use domains
Full article234 words · extracted from arxiv.org · click to collapse
Modern chess engines are silent experts: they play at a superhuman level, but do not offer explanations for their play. On the other hand, language models (LMs) can generate plausible-sounding explanations, but their weak playing strength limits the utility of their explanations. We introduce Queen, a 4B-parameter chess-language model that can explain its moves and plans while playing at the level of a typical Grandmaster. Our novel framework enables domain-specific reasoning through complementary components: an encoder-decoder architecture and an iterative distillation algorithm. This architecture integrates a silent expert chess encoder with an instruction-tuned LM through cross-attention, which we train via a question-answering curriculum to extract chess concepts from the encoder's representations. Building on this domain-adapted model, we iteratively improve its explanations with a natural-language analog of the Bellman update: the model analyzes the positions after its top candidate moves and consolidates them into an explanation of the current position, which is then distilled back into the model. Over seven iterations, our model gains over 900 Elo points (1782 to 2697), substantially surpassing all frontier models on both playing strength and puzzle accuracy, despite containing three orders of magnitude fewer parameters. Furthermore, LM-based evaluations show that our explanations are fluent and approach GPT-5.6-Sol (high) in coherence. The generality of our architecture and training procedure suggests a recipe for applying language models to domains where silent expert encoders are available, like games, robotics, and computer use.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.03695