Three papers report LLM training and distillation gains
Separate papers on diverse self-training, OASIS, and PivotOPD report Qwen3 and Nemotron gains, with conflicting point-versus-percent wording.
The four reports cover three distinct late-September 2026 papers, not a single incident. GROOT treats strategic diversity as the principle for LLM self-training data: hierarchical path sampling beat IID training on hard competitive-programming and next-chapter tasks, and strategically diverse but incorrect traces from Qwen3-4B outperformed IID distillation from a 235B teacher while also seeding reinforcement learning and test-time scaling. OASIS analyzes on-policy self-distillation and finds scaffold correctness matters more than teacher-context correctness; it supervises mostly label-verified on-policy trajectories using only final-answer labels, gaining 3.2 to 3.8 points on average for Qwen3-1.7B, 4B, and 8B across AIME 2024, AIME 2025, and HMMT 2025, and beating standard OPSD by 3.05 points at 8B. PivotOPD, from NVIDIA, trains multi-turn agents with reverse-KL distillation on a teacher gold action and forward-KL on a short recovery sequence after early pivotal mistakes, which both writeups say appear in more than half of failed Qwen3 rollouts from 8B to 235B. Both call it the strongest average of 13 baselines for Qwen3-1.7B and Qwen3-8B on ALFWorld, WebShop, and search-based QA, and both report a Nemotron-3.5 SWE-Bench Verified improvement. Sources disagree on units: summaries say a 5.5-point ALFWorld gain and, in one case, a 3.2-point SWE-Bench gain, while bullets and the later note also phrase those figures as 5.5% (once specified as over the best ALFWorld baseline) and 3.2%.
- GROOT builds self-training data by sampling distinct paths in a hierarchical tree of approaches; adapted Verbalized Sampling instead yields an unstructured set.
- On competitive programming and next-chapter prediction, strategic sampling beat IID training on hard tasks and initialized RL and test-time scaling; diverse incorrect Qwen3-4B traces beat IID distillation from a 235B teacher.
- OASIS keeps the on-policy self-distillation objective but supervises mostly label-verified trajectories and needs only final-answer labels, not written solutions.
- On Qwen3-1.7B, 4B, and 8B, OASIS gained 3.2 to 3.8 points on average across AIME 2024, AIME 2025, and HMMT 2025, and beat OPSD by 3.05 points at 8B.
- Scaffold correctness mattered more than privileged teacher context; unverified scaffolds create an imitation gap that shrinks with scale.
- NVIDIA's PivotOPD found more than half of failed multi-turn rollouts on Qwen3 models from 8B to 235B contain an early pivotal mistake often recoverable in a few guided turns, using reverse-KL on a teacher gold action and forward-KL on…
- PivotOPD was the strongest average among 13 baselines for Qwen3-1.7B and Qwen3-8B on ALFWorld, WebShop, and search-based QA; sources disagree whether the 1.7B ALFWorld result is +5.5 points or +5.5% (one bullet says over the best baseline)…
Coverage timelineoldest first · each row is one article
- · 6d agoStrategically Diverse Sampling for Self-Training
arXiv cs.AI / cs.LG / cs.CL· 52
Strategically diverse sampling for LLM self-training beats IID data, and Qwen3-4B incorrect traces beat a 235B teacher.
- · 3d agoOvercoming Scaling Limits in On-Policy Self-Distillation for LLM Reasoning
Hugging Face daily papers· 55
OASIS keeps on-policy self-distillation effective at scale by supervising verified reasoning trajectories.
- · 2d agoPivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents
Hugging Face daily papers· 52