ROSS: Relearning from Self-Generated Rollouts through Selective Supervision
ROSS reuses historical self-generated rollouts with selective supervision, lifting Qwen3.6 SWE-bench Verified from 64.2% to 68.4%.
ROSS reuses historical self-generated rollouts by keeping the full trajectory as context and applying supervised loss only to selected model continuations. It improves upstream checkpoints across domain-specific reinforcement learning, multi-teacher on-policy distillation, and agentic RL, covering math, code, instruction following, and software engineering. On Qwen3.6-35B-A3B it raises the six-benchmark MOPD average from 58.40% to 62.20% and SWE-bench Verified from 64.20% to 68.40%, using offline SFT without additional policy rollouts.
- Keeps full historical trajectories as context but applies loss only to selected continuations.
- Tested on domain RL, multi-teacher on-policy distillation, and agentic RL.
- On Qwen3.6-35B-A3B, six-benchmark MOPD average rises from 58.40% to 62.20%.
- SWE-bench Verified improves from 64.20% to 68.40% without new policy rollouts.
Full article162 words · extracted from huggingface.co · click to collapse
Large language model post-training generates self-generated rollouts through reinforcement learning and on-policy distillation, yet this experience is often treated as stale once the policy advances. Historical rollouts can remain compatible with a later policy while preserving behaviors that the policy no longer expresses reliably. However, they may also contain mistakes, abandoned attempts, and redundant actions that should not be imitated, motivating finer-grained selective supervision. We introduce ROSS (Relearning from Self-Generated Rollouts through Selective Supervision), which preserves the full historical trajectory as context while applying loss only to selected model-generated continuations. Across domain-specific reinforcement learning, multi-teacher on-policy distillation, and agentic reinforcement learning, ROSS consistently improves upstream checkpoints and outperforms baselines across mathematics, code generation, instruction following, and software engineering. On Qwen3.6-35B-A3B, ROSS improves the six-benchmark MOPD average from 58.40% to 62.20% and SWE-bench Verified from 64.20% to 68.40%. These results show that self-rollout training leaves behind reusable behavioral experience that can yield further gains through offline supervised fine-tuning (SFT), without additional policy rollouts.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.35954