发表机构
Zhejiang University; Ant Group(浙江大学; 蚂蚁集团)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ComputerSD提出一种面向计算机使用智能体的在线自蒸馏方法,利用GUI实时反馈生成引导和步骤级分数,联合优化词元级OPSD与轨迹级GRPO,在OSWorld-Verified上显著超越基线并具备泛化能力。
AI 中文摘要
在线训练使计算机使用智能体(CUA)能够通过与可执行环境的交互来改进自身。然而,现有方法主要依赖稀疏的结果奖励,这无法为中间动作提供监督。同策略自蒸馏(OPSD)通过特权重新评分提供词元级学习信号,但将其直接应用于CUA在线训练面临两个挑战:固定引导可能与学生当前状态失配,且引导引起的概率偏移可能与步骤级正确性冲突。我们提出ComputerSD,一种面向CUA的在线自蒸馏方法,将来自已执行GUI转换的实时反馈转化为策略学习的引导。一个微调的GUI分析器在每个动作后生成引导和步骤级价值分数;引导提供特权上下文,而分数调节产生的OPSD信号。ComputerSD在完全异步训练框架中联合优化词元级OPSD和轨迹级GRPO。在OSWorld-Verified上,ComputerSD在通用Qwen3-VL-8B-Thinking和专用EvoCUA-8B骨干上分别比仅结果GRPO高出1.9和4.1个百分点。分布外设置下的评估进一步支持了ComputerSD的泛化能力。这些结果证明了通过在线自蒸馏从实时反馈中学习对CUA的有效性。
英文摘要
Online training enables computer-use agents (CUAs) to improve through interaction with executable environments. However, existing methods primarily rely on sparse outcome rewards, which provide no supervision for intermediate actions. On-policy self-distillation (OPSD) offers token-level learning signals through privileged rescoring, but directly applying it to CUA online training presents two challenges: fixed guidance may become misaligned with the student's current state, and guidance-induced probability shifts may conflict with step-level correctness. We introduce ComputerSD, an online self-distillation method for CUAs that converts real-time feedback from executed GUI transitions into guidance for policy learning. A fine-tuned GUI analyzer produces guidance and a step-level value score after each action; the guidance provides privileged context, while the score regulates the resulting OPSD signals. ComputerSD jointly optimizes token-level OPSD and trajectory-level GRPO in a fully asynchronous training framework. On OSWorld-Verified, ComputerSD outperforms outcome-only GRPO by 1.9 and 4.1 percentage points on the general-purpose Qwen3-VL-8B-Thinking and specialized EvoCUA-8B backbones, respectively. Evaluation in out-of-distribution settings further supports the generalizability of ComputerSD. These results demonstrate the effectiveness of learning from real-time feedback through online self-distillation for CUAs.
Commentshttps://github.com/ZJU-REAL/ComputerSD