arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UOPD:多轮智能体在线策略蒸馏中的不确定性感知干预

UOPD: Uncertainty-Aware Intervention for On-Policy Distillation of Multi-Turn Agents

Wenbo Zhang, Pengcheng Xu, Weizhi Du, Jing Zhang, Hengrui Cai

arXiv 2609.34036首次发表:更新:

发表机构

University of California, Irvine; University of Michigan, Ann Arbor(加州大学尔湾分校; 密歇根大学安娜堡分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

UOPD通过不确定性感知干预,在关键步骤用教师动作纠正学生,提升多轮智能体在线策略蒸馏的性能,在ALFWorld、WebShop和Search任务上优于现有方法。

AI 中文摘要

在线策略蒸馏(OPD)使用教师提供的密集监督,在学生自身的轨迹上训练学生。在多轮环境中,关键决策步骤上的错误可能会将后续轨迹引向不良结果。我们利用教师对学生动作的低置信度来选择高不确定性步骤进行纠正。在一项受控的ALFWorld研究中,在低置信度步骤上进行一次教师纠正即可改善后续的学生行为和任务成功率,这激励了蒸馏过程中的选择性干预。我们提出了UOPD,一种用于在线策略蒸馏的不确定性感知干预方法。在低不确定性回合中,UOPD执行学生动作并应用标准OPD损失。在高不确定性回合中,它采样并执行教师动作,并通过监督微调训练学生模仿这些动作,这在期望上最小化前向KL散度。UOPD利用自适应不确定性阈值来达到预定的干预率。在实验上,我们在广泛的智能体任务(包括ALFWorld、WebShop和Search)上评估了UOPD,展示了其相对于OPD方法及其变体的优越性能。UOPD相对于标准OPD将WebShop分数提高了高达15.8%。

英文摘要

On-policy distillation (OPD) trains a student on its own rollouts using dense supervision from a teacher. In multi-turn environments, a mistake at a critical decision step can redirect the subsequent rollout toward poor outcomes. We use low teacher confidence on student actions to select high-uncertainty steps for correction. In a controlled ALFWorld study, a single teacher correction at a low-confidence step improves subsequent student behavior and task success, motivating selective intervention during distillation. We propose UOPD, an uncertainty-aware intervention method for on-policy distillation. At low-uncertainty turns, UOPD executes student actions and applies the standard OPD loss. At high-uncertainty turns, it samples and executes teacher actions and trains the student to imitate them through supervised fine-tuning, which minimizes forward Kullback-Leibler divergence in expectation. UOPD utilizes adaptive uncertainty thresholds to target a scheduled intervention rate. Empirically, we evaluate UOPD across a broad range of agentic tasks, including ALFWorld, WebShop, and Search, demonstrating its superior performance over OPD methods and their variants. UOPD improves WebShop score by up to $15.8\%$ relative to standard OPD.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑