arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PivotOPD:在多轮智能体中从关键错误中恢复的学习方法

PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents

Yinghui He, Yapei Chang, Khushi Bhardwaj, Daniele Molinari, Tugrul Konuk, Jan Kautz, Ali Hatamizadeh

arXiv 2609.40285首次发表:更新:

发表机构

Princeton University; NVIDIA; University of Maryland(普林斯顿大学; 英伟达; 马里兰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PivotOPD通过在线策略蒸馏联合训练学生模型预防关键错误并从其状态中恢复,在多个基准上取得最佳平均性能,显著提升任务成功率。

AI 中文摘要

在线策略蒸馏(OPD)是一种训练语言智能体的有前景的方法,它为智能体生成的轨迹提供密集的教师监督。然而,在多轮交互中,一个不正确的动作会改变智能体后续遇到的状态,因此错误会在各轮之间累积。在跨越三个Qwen3模型(8B到235B)的初步实验中,我们发现超过一半的失败轨迹包含一个关键错误,即一个使智能体远离完成任务目标的动作,且该错误通常发生在早期。这些关键错误往往是可以恢复的:在关键轮次之后仅引导模型几轮,就能恢复任务成功。因此,我们提出了PivotOPD,一个在线策略蒸馏框架,它联合训练学生模型以预防关键错误,并从这些错误造成的状态中恢复。在每个关键错误处,教师模型提供一个黄金动作,并在接下来的几轮中为每一轮指定一个恢复动作。预防性蒸馏使用黄金动作和反向KL散度来引导学生远离关键错误,而恢复性蒸馏使用恢复动作和前向KL散度来转移学生很少采样的恢复行为。在ALFWorld、WebShop和基于搜索的问答上,与13个基线相比,PivotOPD在Qwen3-1.7B和Qwen3-8B学生模型上均取得了最强的平均性能,其中1.7B学生模型在ALFWorld上比最强基线提高了+5.5%。这些收益也转移到了软件工程领域的另一个模型家族,PivotOPD将Nemotron-3.5学生在SWE-Bench Verified上的解决率提高了+3.2%。项目页面:此HTTPS URL。

英文摘要

On-policy distillation (OPD) is a promising approach for training language agents, providing dense teacher supervision on student-generated trajectories. However, in multi-turn interaction, an incorrect action changes the states the student encounters later, so errors compound across turns. In preliminary experiments across three Qwen3 models (8B to 235B), we find that more than half of the failed rollouts contain a pivotal mistake, an action that moves the agent farther from completing the task, and this mistake typically occurs early. These pivotal mistakes often remain recoverable: guiding the model for only a few turns after the pivotal turn can restore task success. We therefore propose PivotOPD, an on-policy distillation framework that jointly trains the student to prevent pivotal mistakes and to recover from the states they create. At each pivotal mistake, a teacher model provides a gold action and then names a recovery action at each of the next few turns. Preventive distillation uses the gold action with reverse KL to steer the student away from the pivotal mistake, while recovery distillation uses the recovery actions with forward KL to transfer recovery behaviors that the student rarely samples. Against 13 baselines on ALFWorld, WebShop, and Search-based QA, PivotOPD achieves the strongest average performance for both Qwen3-1.7B and Qwen3-8B students, improving over the strongest baseline on ALFWorld by +5.5% with the 1.7B student. The gains also transfer to another model family on the software engineering domain, where PivotOPD raises the resolve rate of a Nemotron-3.5 student on SWE-Bench Verified by +3.2%. Project page: https://research.nvidia.com/labs/lpr/pivotopd/

CommentsPivotOPD technical report; Project page: https://research.nvidia.com/labs/lpr/pivotopd/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑