AdviSD: Learning to Advise Frontier LLMs via Targeted Multi-Turn Self-Distillation
AdviSD trains a small advisor with selective self-distillation, beating advisor-GRPO on BFCL-v3 and EnvScaler.
AdviSD trains a small advisor to steer a frozen language-model executor with natural-language advice. It combines outcome-based reinforcement learning with selective self-distillation: reflection proposes corrections, and only decisions whose advice meaningfully changes the scored executor response are used for supervision, without executor likelihoods or extra rollouts. Qwen3-8B advisors for Gemini and Claude beat advisor-GRPO by 4.2 to 6.4 percentage points on BFCL-v3 and by 3.9 to 5.1 score points on EnvScaler. The advisors generalize out of domain and transfer across executor versions and model families.
- A small advisor steers a frozen executor using natural-language advice.
- AdviSD pairs outcome reinforcement learning with selective self-distillation from scored corrections.
- Selection uses how much issued advice changes the score of a recorded response.
- Qwen3-8B advisors gain 4.2 to 6.4 points on BFCL-v3 over advisor-GRPO.
- Trained advisors transfer across out-of-domain tasks, executor versions, and model families.
Full article215 words · extracted from arxiv.org · click to collapse
A small trainable advisor can steer a frozen language-model executor using natural-language advice. In addition to learning from task rewards, the advisor can use feedback from completed interactions to improve its advice. However, a plausible correction need not change execution, yet learning from such corrections can still affect the advisor's future decisions in other contexts. In a shared-parameter model, we prove that such corrections can limit learning if their targets favor useful advice less strongly than those of other corrections. Keeping them less often than the rest improves the model's eventual performance compared to learning from every correction. Motivated by this, our method, Advisor Self-Distillation (AdviSD), pairs outcome-based reinforcement learning with self-distillation from a feedback-conditioned copy of the advisor selectively. Reflection proposes corrections, and the advisor scores the same recorded executor response with and without its issued advice, using the magnitude of the difference to select decisions for supervision. This approach does not require executor likelihoods or additional executor rollouts. Experiments with Qwen3-8B advisors for Gemini and Claude show that AdviSD outperforms advisor-GRPO by 4.2-6.4 percentage points on BFCL-v3 and by 3.9-5.1 score points on EnvScaler. The trained advisors generalize to out-of-domain tasks and transfer across different executor versions and model families. AdviSD also beats matched-count random selection, supporting the value of its selection rule.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.38142