发表机构
University of Chicago; Meituan LongCat Interaction Team; University of the Chinese Academy of Sciences; The Hong Kong University of Science and Technology (Guangzhou)(芝加哥大学; 美团龙猫交互团队; 中国科学院大学; 香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
DivOPD通过分散回合预算并优先高分歧回合,解决异步多轮在策略蒸馏中的批次过时问题,在多个基准上提升成功率并加速训练。
AI 中文摘要
在策略蒸馏(OPD)通过教师对学生在环境中的交互进行监督来训练学生智能体。然而,在异步多轮训练中,到达顺序的批处理可能使少数早期或较长的轨迹主导学习者更新,而其他有效轨迹在使用前就已过时,从而浪费已生成的体验。为解决此问题,我们提出DivOPD,一种简单的学习者侧批次选择方法,它将固定的回合预算分散到更多轨迹上,并在每个轨迹内优先处理教师-学生累积分歧较大的回合。没有可用教师反馈的回合被排除在外。每回合的损失和优化器保持不变;选择仅改变哪些学生访问的回合获得训练权重。对于无进展的轨迹,一个可选扩展会短暂地将控制权交给教师,然后再交还给学生。在模拟的ALFWorld、ScienceWorld和WebShop基准上的六种教师-学生设置中,使用1.5B-7B学生,DivOPD将跨设置的平均峰值成功率从77.4提高到84.4,并将最后五次评估的平均成功率从71.5提高到78.6。它达到了所有报告的特定设置目标,相对于普通OPD,训练令牌的几何平均加速为1.84倍,学习者GPU时间为1.87倍。教师干预进一步将最后五次平均值提高到82.4,同时相对于普通OPD保持约1.7倍的学习者GPU加速。代码将在该https URL发布。
英文摘要
On-policy distillation (OPD) trains student agents through teacher supervision on their own interactions with an environment. However, in asynchronous multi-turn training, arrival-order batching can allow a few early or long rollouts to dominate learner updates while other valid rollouts become stale before being used, wasting already-generated experience. To address this problem, we introduce DivOPD, a simple learner-side batch-selection method that spreads a fixed turn budget across more rollouts and, within each rollout, prioritizes turns with larger cumulative teacher-student disagreement. Turns without usable teacher feedback are excluded. The per-turn loss and optimizer remain fixed; selection only changes which student-visited turns receive training weight. For no-progress rollouts, an optional extension briefly hands control to the teacher before returning it to the student. Across six teacher-student settings on the simulated ALFWorld, ScienceWorld, and WebShop benchmarks, with 1.5B-7B students, DivOPD raises cross-setting mean peak success rate from 77.4 to 84.4 and mean success over the last five evaluations from 71.5 to 78.6. It reaches all reported setting-specific targets with geometric-mean speedups of 1.84x in training tokens and 1.87x in learner GPU time relative to vanilla OPD. Teacher intervention further raises this last-five mean to 82.4 while retaining about 1.7x learner-GPU speedup over vanilla OPD. Code will be released at https://github.com/HanyangWang0418-oss/DivOPD.
Comments24 pages, 9 figures, 19 tables. Code: https://github.com/HanyangWang0418-oss/DivOPD