PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents
PivotOPD trains agents to avoid and recover from early pivotal mistakes, lifting ALFWorld and SWE-Bench scores.
PivotOPD is an on-policy distillation method for multi-turn language agents that both avoids pivotal mistakes and recovers from the states they create. Preliminary runs on Qwen3 models from 8B to 235B found that more than half of failed rollouts contain an early pivotal mistake, yet a few teacher-guided turns can restore success. Against 13 baselines on ALFWorld, WebShop, and search-based QA, it is the strongest average method for Qwen3-1.7B and Qwen3-8B students, including a 5.5-point ALFWorld gain for the 1.7B student. The approach also raises a Nemotron-3.5 student's SWE-Bench Verified resolve rate by 3.2%.