Self-Retrospection Distillation: Turning Post-hoc Experiences into Prior Foresight
Self-Retrospection Distillation turns post-hoc agent trajectories into foresight, helping when rollout rewards lack contrast.
The paper introduces prospective learning and Self-Retrospection Distillation (SRD), which uses completed trajectories to supervise foresight predictions from the pre-interaction view. Foresight is only a training target and need not be generated at inference. Across 10 tool-integrated reasoning and long-horizon agentic tasks, SRD adds up to 24.2 percentage points over RLVR and self-distillation baselines. When 37–98% of rollout groups share the same reward, SRD still extracts signal; in a 2B setting where 98% of groups were all failures, RLVR stayed at 0.0% success while SRD reached 60.6% under the same rollout budget.
- SRD distills hindsight from finished trajectories into trajectory-blind foresight.
- Foresight is a training target and need not be generated at inference.
- Gains reach 24.2 percentage points across 10 agentic and tool-use tasks.
- In a 2B all-failure setting, SRD raised success from 0.0% to 60.6%.
Full article240 words · extracted from huggingface.co · click to collapse
Reinforcement learning with verifiable rewards (RLVR) turns agent experience into learning signals primarily through scalar outcome rewards after interaction. For group-relative objectives, however, this signal vanishes when all rollouts receive the same reward, even though their trajectories may reveal useful information about what the task requires and how the agent fails. We ask a complementary question: can hindsight teach an agent what it could have anticipated before acting? We introduce prospective learning, which uses post-hoc experience to supervise foresight predictions from the pre-interaction view, and instantiate it with Self-Retrospection Distillation (SRD). Intuitively, a completed trajectory reveals knowledge that would have been useful and pitfalls that should be avoided; SRD distills this privileged hindsight into trajectory-blind foresight of the same policy. Foresight serves only as a training target and need not be explicitly generated at inference time. Across 10 tool-integrated reasoning and long-horizon agentic tasks, SRD complements RLVR and self-distillation baselines with gains of up to 24.2 pp. Its advantage is especially pronounced when reward contrast is scarce: when 37--98% of rollout groups are reward-uniform across model scales, yet SRD can still exploit learning signal from sampled trajectories. In the 2B setting, where 98% of groups are all-failure, the RLVR training ends up at 0.0% success, while adding SRD reaches 60.6% under the same rollout budget. Our results suggest that post-hoc agent experience is useful not only for evaluating or improving behavior, but also for shaping predictive representations before available interaction.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.08077