AI 中文总结
本文针对移动GUI智能体训练中丢失动作依据的问题,提出门控事后蒸馏方法,利用下一张截图作为特权信息,在AndroidWorld等基准上相比GRPO提升了任务成功率。
AI 中文摘要
GUI智能体通常从成功的交互轨迹中进行离线训练,标准训练将每条轨迹分解为前缀-动作对:智能体根据当前屏幕和交互历史预测动作,而后续观测会被丢弃,这就丢失了动作正确的依据,因为证据往往仅出现在后续屏幕中。例如,要启用Soft Wrap,智能体应点击“编辑”或“视图”,但菜单打开前不会显示该选项;没有这类证据,标准模仿学习几乎无法让模型采样并学习正确的推理。为解决此问题,我们提出门控事后蒸馏(GHD),该方法在训练时将下一张截图作为特权信息:学生从可观测的轨迹前缀进行预测,而参数共享的教师额外观测下一张截图,并重新评分学生的在线策略响应;仅当学生失败且事后条件教师恢复了演示动作时,才应用蒸馏。我们在AndroidWorld和AndroidLab上针对两个视觉语言模型开展实验,结果表明,该方法相比GRPO提升了任务成功率,代码和检查点将公开提供。
英文摘要
GUI agents are commonly trained offline from successful interaction trajectories. Standard training decomposes each trajectory into prefix-action pairs: the agent predicts an action from the current screen and interaction history, while the subsequent observation is discarded. This removes the rationale of why an action is correct: the evidence often appears only on the subsequent screen. For example, to enable Soft Wrap, the agent should click Edit or View, but nothing reveals this until the menu opens. Without such evidence, standard imitation gives the model little chance of ever sampling and thus learning the correct reasoning. To address this issue, we propose Gated Hindsight Distillation (GHD), which uses the next screenshot as privileged information during training. A student predicts from the observable trajectory prefix, while a parameter-sharing teacher additionally observes the next screenshot and re-scores the student's on-policy responses. We apply distillation only when the student fails and the hindsight-conditioned teacher recovers the demonstrated action. GHD improves task success over GRPO on AndroidWorld and AndroidLab across two vision-language models. The code and checkpoints will be made available.