Three on-policy distillation papers report Qwen3 gains
PivotOPD, Neighborhood OPSD, and a multi-teacher study report Qwen3 agent, math, and optimizer effects, with sources split on points versus percents.
Three related on-policy distillation papers were reported on Hugging Face daily papers and arXiv from 29 September through 1 October 2026. PivotOPD trains multi-turn agents to avoid early pivotal mistakes with reverse-KL distillation on a teacher gold action and to recover with forward-KL on later actions; across Qwen3 models from 8B to 235B, more than half of failed rollouts had such a mistake, often recoverable in a few guided turns. It was the strongest average among 13 baselines for Qwen3-1.7B and Qwen3-8B on ALFWorld, WebShop, and search-based QA, but sources disagree on whether the 1.7B ALFWorld gain is 5.5 points or 5.5% and whether it is over the best baseline, and they likewise give the Nemotron-3.5 SWE-Bench Verified gain as 3.2 points or 3.2%. Neighborhood OPSD builds a compact pool of frozen experts and uses MaxPeak plus quantile routing for a clipped forward-KL target, lifting Average@12 over standard OPSD by 2.75, 1.67, and 1.94 points on Qwen3-1.7B, 4B, and 8B on AIME 2024, AIME 2025, and HMMT February 2025. A separate Qwen3-1.7B multi-teacher study finds averaging favors longer responses—called loss averaging in one report and token averaging in another—while Adam raises update similarity from 0.83 between teachers to 0.96 between averaging rules, BF16 shows far fewer changed weights than FP32 (7-11% versus about 97%), and math accuracy moves by 2.6 or 2.1 points depending on the averaging rule.
- PivotOPD (attributed to NVIDIA in one report) uses reverse-KL distillation on a teacher gold action and forward-KL distillation on later recovery actions.
- On Qwen3 models from 8B to 235B, more than half of failed multi-turn rollouts contain an early pivotal mistake that is often recoverable within a few guided turns.
- PivotOPD has the strongest average among 13 baselines for Qwen3-1.7B and Qwen3-8B on ALFWorld, WebShop, and search-based QA; sources give the 1.7B ALFWorld result as 5.5 points or 5.5%, and one says it is over the best baseline.
- A Nemotron-3.5 student's SWE-Bench Verified resolve rate is reported up 3.2 points in one account and 3.2% in another.
- Neighborhood OPSD raises Average@12 over standard OPSD by 2.75, 1.67, and 1.94 points on Qwen3-1.7B, 4B, and 8B across AIME 2024, AIME 2025, and HMMT February 2025; inference uses only the student.
- A multi-teacher study uses Qwen3-1.7B and four RL domain teachers from the same initialization, plus SmolLM3-3B diagnostics; Adam's first moment raises update cosine similarity from 0.83 between teachers to 0.96 between averaging rules.
- About 97% of FP32 master weights differ from initialization versus 7-11% of BF16 weights; one report says math accuracy is 2.6 points higher or 2.1 points lower by averaging rule, while another says it shifts by 2.6 or 2.1 points.
- Sources disagree on whether response-length bias comes from loss averaging or token averaging.
Coverage timelineoldest first · each row is one article
- · 9d agoPivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents
Hugging Face daily papers· 52
PivotOPD trains multi-turn agents to avoid pivotal mistakes and recover within a few turns.
- · 9d agoBetter Supervision Is Nearby: Neighborhood On-Policy Self-Distillation
Hugging Face daily papers· 52
Neighborhood OPSD lifts Qwen3 math scores by routing student prefixes to nearby expert teachers.