Shockingly Simple Self-retrospection Improves Agentic Models Without RL
Retrospection-only fine-tuning lifts Qwen3.5-4B on SWE-bench without rewards or an external teacher.
Retrospection-Only Fine-Tuning (ROFT) has an agent attempt a task, write a retrospective explanation, and train with next-token prediction on those explanation tokens alone. Using Qwen3.5-4B on mixed successful and failed attempts, ROFT reached 49.2% on SWE-bench Verified and 26.8% on SWE-bench Pro after 20 updates without a verifier, versus GRPO at 48.0% and 25.3% after 40 updates. It also solved tasks where all 64 base-model samples failed, and prompting for more direct solutions shortened later attempts without a length penalty.
- ROFT fine-tunes only on self-generated retrospective explanation tokens.
- No external teacher, verifier, or reward-based policy update is used.
- Qwen3.5-4B reaches 49.2% SWE-bench Verified and 26.8% Pro after 20 updates.
- It solves tasks where all 64 sampled base-model attempts failed.
Full article250 words · extracted from arxiv.org · click to collapse
People learn not only by repeating successful actions, but also by recounting and explaining their experiences, revising their understanding to guide future behavior. Can a language-model agent improve its future actions by training only on explanations of its own experience? We investigate this question by studying Retrospection-Only Fine-Tuning (ROFT), a minimal online procedure designed to isolate the effect of explanation-only training on subsequent behavior. The agent attempts a task, observes available feedback, generates a retrospective explanation, and is fine-tuned with a next-token prediction loss on the explanation tokens alone. The procedure uses neither an external teacher nor a reward-based policy update. In software-engineering experiments with Qwen3.5-4B, ROFT is trained on problems with mixed successful and unsuccessful base-model attempts. On held-out SWE-bench Verified and Pro, it reaches 49.2% and 26.8% solve rates after 20 updates without using a verifier, compared with GRPO's 48.0% and 25.3% after 40 updates in the evaluated runs, and makes faster early progress in training time and sampled attempts. It also learns to solve individual tasks on which all 64 sampled base-model attempts failed, showing that learning can begin without any initially successful trajectories. Behavioral analyses find that ROFT indirectly assigns credit to actions, encouraging good actions and discouraging incorrect ones. Moreover, prompting retrospections to emphasize more direct solutions yields shorter subsequent attempts even without an explicit length penalty. Together, these findings show that learning to explain can also improve learning to do, establishing self-generated retrospections as useful training targets and motivating further study of explanation-to-action transfer.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.35741