PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents
PivotOPD trains multi-turn agents to avoid pivotal mistakes and recover within a few turns.
Across three Qwen3 models from 8B to 235B, more than half of failed multi-turn rollouts contain an early pivotal mistake that is often recoverable within a few guided turns. PivotOPD, from NVIDIA, uses reverse-KL distillation on a teacher gold action to prevent that mistake and forward-KL distillation on subsequent recovery actions. It records the strongest average among 13 baselines for Qwen3-1.7B and Qwen3-8B on ALFWorld, WebShop, and search-based QA, including a 5.5-point ALFWorld gain for the 1.7B student. A Nemotron-3.5 student also improves SWE-Bench Verified resolve rate by 3.2 points.