arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35319cs.LGcs.AI

师生差距并不足够:面向多轮自主智能体的结果引导在线蒸馏

Teacher-Student Gaps Are Not Enough: Outcome-Guided On-Policy Distillation for Multi-Turn Autonomous Agents

Tong Zhang, Zhou Liu, Yihao Liu, Jiahua Bao, Xuchen Li, Honglin Lin, Tao Cheng, Zhihan Yu, Kai Tang, Xiaoxi Jiang, Guanjun Jiang

首次发表
浏览论文内容

中文总结 AI 辅助

针对多轮自主智能体,提出结果引导在线蒸馏(OG-OPD),利用最终任务结果校准监督权重,解决师生差距与收益不匹配问题,在多个基准上显著提升任务成功率。

中文摘要 AI 辅助

在线蒸馏(OPD)在学生的自身轨迹上训练学生模型,并辅以密集的教师监督。近期关于多轮自主智能体的在线蒸馏研究往往将较大的师生词元级分布差距视为有前景的干预点,将更大的差距与更大的修正需求联系起来。然而,我们的实证分析揭示了一种监督收益不匹配现象:大的差距可能是良性的,而小的差距可能对结果至关重要。师生差距捕捉的是当前回合的差异,而教师指导的收益取决于当前学生随后如何与环境交互。学生即使选择了与教师不同的动作,仍可能成功完成任务;而教师偏好的动作可能导致学生无法完成任务的后续状态。因此,仅凭局部差距不足以判断教师指导是否对当前学生有益。有效的监督应侧重于当前学生能够转化为更好最终任务结果的指导。据此,我们提出了结果引导在线蒸馏(OG-OPD),该方法对教师监督应用轨迹相对权重,并利用配对的学生后续轨迹的最终任务结果来校准这些权重。这种校准在教师指导对当前学生有益的回合中,选择性地加强学生原始轨迹上的监督。在ALFWorld、ScienceWorld和WebShop上,OG-OPD在各种设置下均持续优于基线方法。与原始OPD相比,它将任务成功率提高了3.6至17.7个百分点,相较于最强基线最高提升了7.0个百分点。

英文摘要

On-policy distillation (OPD) trains a student on its own trajectories with dense teacher supervision. Recent work on OPD for multi-turn autonomous agents often treats large teacher-student token-level distributional gaps as promising intervention points, linking larger gaps to a greater need for correction. Yet, our empirical analysis reveals a supervision-benefit mismatch: large gaps can be benign, while small gaps can be outcome-critical. Teacher-student gaps capture differences at the current turn, whereas the benefit of teacher guidance depends on how the current student interacts with the environment afterward. The student may still succeed despite choosing an action that differs from the teacher's, while a teacher-preferred action may lead to a state from which the student cannot complete the task. Local gaps alone are therefore not enough to determine whether teacher guidance benefits the current student. Effective supervision should instead emphasize guidance that the current student can translate into better final task outcomes. Accordingly, we propose Outcome-Guided On-Policy Distillation (OG-OPD), which applies trajectory-relative weighting to teacher supervision and calibrates these weights using final task outcomes from paired student continuations. This calibration selectively strengthens supervision on the student's original trajectories at turns where teacher guidance benefits the current student. Across ALFWorld, ScienceWorld, and WebShop, OG-OPD consistently outperforms baselines under diverse settings. It improves task success rates by 3.6-17.7 percentage points over vanilla OPD and by up to 7.0 percentage points over the strongest baseline.

发表机构

  • Qwen Large Model Application Team, Alibaba(阿里巴巴通义千问大模型应用团队)
  • Peking University(北京大学)
  • University of Chinese Academy of Sciences(中国科学院大学)
  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

↑