Self-Play Search Distillation for Large Language Model Reasoning
Self-play search distillation raises Qwen3-4B-Base's math benchmark mean from 24.1 to 36.6.
Self-Play Search Distillation generates synthetic reasoning data from MuZero-like networks trained by self-play on board games. At each state the search record names a preferred decision, plausible alternatives, opponent replies, and value estimates, which are converted into environment-grounded chains of thought. Although trained only on those game records, Qwen3-4B-Base raises its mean over six mathematics benchmarks from 24.1 to 36.6 and its held-out-game win rate from 15% to 45%. The authors present SPSD as an annotation-efficient alternative to low-quality synthetic data and costly human labels.
- SPSD turns MuZero-like self-play search into chain-of-thought supervision.
- Qwen3-4B-Base mean score on six math benchmarks rises from 24.1 to 36.6.
- Held-out game win rate increases from 15% to 45%.
- Training uses only game search records yet transfers to unseen mathematics.
Full article155 words · extracted from huggingface.co · click to collapse
Improving reasoning abilities in Large Language Models (LLMs) requires high-quality data that exposes difficult decisions, competing alternatives, and their consequences. Data scarcity is driven by the low quality of synthetic data and the cost of human labeling. We introduce Self-Play Search Distillation (SPSD), a framework for generating superhuman synthetic data via self-play of MuZero-like networks trained on board games. SPSD uses executable environments to turn search into structured reasoning problems. At each state, the expert identifies a preferred decision, plausible alternatives, plausible opponent replies, and value estimates. By converting the self-play search records into superhuman chains-of-thought, we train LLMs with environment-grounded supervision. Although trained only on self-play search records, SPSD transfers to unseen mathematics. On Qwen3-4B-Base, it raises the mean over six mathematics benchmarks from 24.1 to 36.6 while increasing the held-out-game win rate from 15% to 45%. SPSD offers an annotation-efficient way to create high-quality synthetic data for improving LLM performance in reasoning tasks.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.30936