ComputerSD: Online Self-Distillation from Real-Time Feedback for Computer-Use Agents
ComputerSD turns real-time GUI feedback into online self-distillation signals for computer-use agents.
ComputerSD is an online self-distillation method for computer-use agents that converts executed GUI transitions into policy guidance. A fine-tuned GUI analyzer emits guidance and a step-level value score that regulates on-policy self-distillation, trained jointly with trajectory-level GRPO in an asynchronous framework. On OSWorld-Verified it beat outcome-only GRPO by 1.9 points on Qwen3-VL-8B-Thinking and 4.1 points on EvoCUA-8B, with further out-of-distribution gains.
- GUI analyzer supplies guidance and step-level value scores.
- Jointly optimizes token-level OPSD and trajectory-level GRPO.
- Gains of 1.9 and 4.1 points on OSWorld-Verified.
- Tested on Qwen3-VL-8B-Thinking and EvoCUA-8B.
Full article178 words · extracted from arxiv.org · click to collapse
Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers token-level learning signals through privileged rescoring, but directly applying it to CUA online training presents two challenges: fixed guidance may become misaligned with the student's current state, and guidance-induced probability shifts may conflict with step-level correctness. We introduce ComputerSD, an online self-distillation method for CUAs that converts real-time feedback from executed GUI transitions into guidance for policy learning. A fine-tuned GUI analyzer produces guidance and a step-level value score after each action; the guidance provides privileged context, while the score regulates the resulting OPSD signals. ComputerSD jointly optimizes token-level OPSD and trajectory-level GRPO in a fully asynchronous training framework. On OSWorld-Verified, ComputerSD outperforms outcome-only GRPO by 1.9 and 4.1 percentage points on the general-purpose Qwen3-VL-8B-Thinking and specialized EvoCUA-8B backbones, respectively. Evaluation in out-of-distribution settings further supports the generalizability of ComputerSD. These results demonstrate the effectiveness of learning from real-time feedback through online self-distillation for CUAs.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.40253