Selecting Diverse SFT Traces Improves Post-RL Generalization
Selecting diverse reasoning routes for SFT improves later reinforcement-learning generalization on math and puzzles.
The study finds verified solutions are not equally useful for preparing reasoning models for reinforcement learning and proposes a lightweight rule-based fingerprint to select diverse reasoning routes in SFT data. From one pool and budget, diverse rather than similar routes improve post-RL coverage on puzzles and mathematics, including harder problems. Route-diverse SFT raises OLMo3-7B pass@8 by 16.9 points on environments held out from SFT. In a single-model setting, diverse selection gains up to 6.2 points of mean pass@8 across 10 math benchmarks, and a CPU-only selector beats costlier alternatives on three open-source corpora.
- A rule-based fingerprint selects diverse reasoning routes without model calls.
- Diverse SFT lifts OLMo3-7B pass@8 by 16.9 points on held-out environments.
- Single-model diverse selection gains up to 6.2 mean pass@8 on 10 math benchmarks.
- Diverse traces give group-relative RL more prompts with a learning signal.
Full article193 words · extracted from huggingface.co · click to collapse
Verified solutions are not equally useful for preparing reasoning models for reinforcement learning (RL). We present a comprehensive study of route diversity, the variation in the sequences of reasoning steps in supervised fine-tuning (SFT) data, and propose a lightweight, rule-based fingerprint to select for it. From one pool at one budget, with matched training recipes and checkpoints, selecting diverse rather than similar routes improves post-RL problem coverage across puzzles and mathematics, including on problems harder than those seen in either training stage. In synthetic experiments, route-diverse SFT improves OLMo3-7B's pass@8 by 16.9 points on environments held out from SFT. In a single-model condition, where one model writes every candidate, diverse selection gains up to 6.2 points of mean pass@8 across 10 mathematics benchmarks. Pre-RL diagnostics suggest why: diverse SFT can produce both successful and failed attempts on more prompts despite slightly lower mean accuracy, giving group-relative RL more prompts with a learning signal. On 3 open-source corpora, our CPU-only selector, without model calls, outperforms more expensive alternatives in every comparison of mean post-RL performance. These results identify reasoning-route diversity as a practical criterion for selecting SFT data that better prepares models for RL.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.33780