发表机构
Fudan University; Tencent Hunyuan(复旦大学; 腾讯文言)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对在线策略蒸馏在长期智能体任务应用不足的问题,提出TurnOPD,通过自适应展开深度预算和渐进式转向归一化损失预算两个策略,在多任务实验中实现更高验证准确率,超越普通OPD,推进了准确率-时间前沿。
AI 中文摘要
在线策略蒸馏(OPD)通过在学生自身轨迹上与更强的教师进行匹配来训练学生策略,为语言智能体训练提供了一个有前景的框架。但其在长期智能体任务中的应用尚未得到充分探索。我们指出了普通智能体OPD的两个关键低效问题:全视野展开在尾部转向时浪费资源且提供的KL监督微弱且有噪声,轨迹级KL目标将大部分损失集中在浅层 tokens 上,导致深层决策转向训练不足。为解决这些挑战,我们提出了TurnOPD,一种用于长期智能体高效在线策略蒸馏的转向级预算策略。TurnOPD由两个预算控制器组成:自适应展开深度预算,使用基于探测的转向统计来确定展开长度;渐进式转向归一化损失预算,逐渐将KL加权从token级转移到转向平衡监督。在ALFWorld、WebShop和Multi-Hop Search上使用任务专用教师模型进行的实验表明,TurnOPD在相同的训练预算下实现了更高的验证准确率,并超越了普通OPD的准确率-时间前沿。
英文摘要
On-policy distillation (OPD) trains a student policy by matching a stronger teacher on the student's own trajectories, offering a promising framework for language agent training. However, its application to long-horizon agentic tasks remains insufficiently explored. We identify two key inefficiencies in vanilla agent OPD: (1) full-horizon rollouts often waste wall-clock resources on tail turns that provide weak and noisy KL supervision, and (2) trajectory-level KL objectives concentrate most of the loss on shallow tokens, leaving deeper decision turns under-trained once initial behaviors are aligned. To address these challenges, we propose TurnOPD, a turn-level budgeting strategy for efficient on-policy distillation of long-horizon agents. TurnOPD consists of two budget controllers: adaptive rollout-depth budgeting, which uses probe-based turn statistics to determine rollout length, and progressive turn-normalized loss budgeting, which gradually shifts KL weighting from token-level to turn-balanced supervision. Experiments on ALFWorld, WebShop, and Multi-Hop Search with task-specialized teacher models show that TurnOPD achieves superior validation accuracy under equal wall-clock training budgets and advances the accuracy--time frontier beyond vanilla OPD.