arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TurnOPD:使在线策略蒸馏具有转向意识以实现高效的长期智能体训练

TurnOPD: Making On-Policy Distillation Turn-Aware for Efficient Long-Horizon Agent Training

Yuhang Zhou, Kai Zheng, Haoling Li, Dengyun Peng, Can Xu, Jingjing Chen

arXiv 2607.05804首次发表:更新:

发表机构

Fudan University; Tencent Hunyuan(复旦大学; 腾讯文言)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对在线策略蒸馏在长期智能体任务应用不足的问题,提出TurnOPD,通过自适应展开深度预算和渐进式转向归一化损失预算两个策略,在多任务实验中实现更高验证准确率,超越普通OPD,推进了准确率-时间前沿。

AI 中文摘要

在线策略蒸馏(OPD)通过在学生自身轨迹上与更强的教师进行匹配来训练学生策略,为语言智能体训练提供了一个有前景的框架。但其在长期智能体任务中的应用尚未得到充分探索。我们指出了普通智能体OPD的两个关键低效问题:全视野展开在尾部转向时浪费资源且提供的KL监督微弱且有噪声,轨迹级KL目标将大部分损失集中在浅层 tokens 上,导致深层决策转向训练不足。为解决这些挑战,我们提出了TurnOPD,一种用于长期智能体高效在线策略蒸馏的转向级预算策略。TurnOPD由两个预算控制器组成:自适应展开深度预算,使用基于探测的转向统计来确定展开长度;渐进式转向归一化损失预算,逐渐将KL加权从token级转移到转向平衡监督。在ALFWorld、WebShop和Multi-Hop Search上使用任务专用教师模型进行的实验表明,TurnOPD在相同的训练预算下实现了更高的验证准确率,并超越了普通OPD的准确率-时间前沿。

英文摘要

On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories, offering a promising framework for language agent training. However, its application to long-horizon agentic tasks remains insufficiently explored. We identify two key inefficiencies in vanilla agent OPD: (1) full-horizon rollouts often waste wall-clock resources on tail turns that provide weak and noisy KL supervision, and (2) trajectory-level KL objectives concentrate most of the loss on shallow tokens, leaving deeper decision turns under-trained once initial behaviors are aligned. To address these challenges, we propose TurnOPD, a turn-level budgeting strategy for efficient on-policy distillation of long-horizon agents. TurnOPD consists of two budget controllers: adaptive rollout-depth budgeting, which uses probe-based turn statistics to determine rollout length, and progressive turn-normalized loss budgeting, which gradually shifts KL weighting from token-level to turn-balanced supervision. Experiments on ALFWorld, WebShop, and Multi-Hop Search with task-specialized teacher models show that TurnOPD achieves superior validation accuracy under equal wall-clock training budgets and advances the accuracy--time frontier beyond vanilla OPD.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑