Strategically Diverse Sampling for Self-Training
Strategically diverse sampling for LLM self-training beats IID data, and Qwen3-4B incorrect traces beat a 235B teacher.
The paper treats strategic diversity, not only correctness, as the principle for building LLM self-training data. GROOT constructs a hierarchical tree of approaches and samples distinct paths, while adapted Verbalized Sampling produces an unstructured set of approaches. On competitive programming and next-chapter prediction, strategically sampled training beat IID training on hard tasks and initialized reinforcement learning and test-time scaling. Self-training on strategically diverse but incorrect traces from Qwen3-4B outperformed IID distillation from a 235B teacher.
- GROOT samples distinct paths through a hierarchical tree of approaches.
- Models trained on strategic data beat IID-trained models on hard tasks.
- Diverse incorrect Qwen3-4B traces beat IID distillation from a 235B teacher.
- Strategic data gives strong starts for RL and test-time scaling.
Full article176 words · extracted from arxiv.org · click to collapse
Many LLM training and inference methods, including RL and test-time scaling, depend on repeated sampling, but benefit only when the responses meaningfully differ. Self-training faces the same challenge: training data is typically constructed by sampling IID responses and filtering primarily for correctness, thereby overrepresenting strategies a model already favours. We investigate strategic diversity, or substantive variation among approaches to a problem, as an alternative principle for constructing self-training data. We generate strategically diverse data with two sampling methods: GROOT, a new method which constructs a hierarchical tree of approaches and samples distinct paths, and Verbalized Sampling (VS), adapted to produce an unstructured set of approaches. Across competitive programming and Next-Chapter Prediction domains, models trained on strategically sampled data outperform IID-trained counterparts on difficult tasks and provide strong initializations for RL and test-time scaling. Most strikingly, self-training on strategically diverse but incorrect traces from Qwen3-4B outperforms IID distillation from a 235B teacher. These results challenge prevailing assumptions about what makes useful self-training data and show that diversity of approaches can matter more than correctness or teacher scale.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.31571