AI 中文总结
该研究针对LLM智能体现有自蒸馏方法的局限,提出OVCSD方法,经实验在ALFWorld和WebShop数据集上显著提升了智能体任务完成成功率。
AI 中文摘要
近期关于大语言模型(LLM)智能体的研究正从外部能力引导转向能力内化,使智能体在推理时无需检索即可保留有用技能。在线策略自蒸馏(OPSD)提供了一个有前景的方向,但现有许多方法通常通过对智能体生成轨迹中的行动进行评分来监督学生模型。这种监督存在两个局限:教师偏好未通过环境结果验证,且行动级评分未充分利用智能体探索、教师探索及其行为关系中的信息。因此,我们提倡结果验证式教师监督及师生轨迹间的对比学习。基于此,我们提出结果验证式对比自蒸馏(OVCSD)。OVCSD将失败的智能体探索组织成前缀树,从智能体到达的状态自适应调用技能条件教师,仅保留经结果验证的成功延续。随后在首个状态对齐分歧处应用局部对比学习,蒸馏分歧后的教师后缀以传递完成行为。在ALFWorld和WebShop上针对三种模型规模的实验表明,OVCSD始终优于无技能强化学习(RL)及现有自蒸馏基线,在ALFWorld和WebShop上较最强基线分别实现了29.7和5.4的绝对成功率提升,且训练期间添加的特权交互占比不足3%。
英文摘要
Recent work on LLM agents is shifting from external capability elicitation to capability internalization, enabling agents to retain useful skills without retrieval at inference time. On-policy self-distillation (OPSD) offers a promising direction, but many existing methods typically supervise students by scoring actions along student-generated trajectories. Such supervision has two limitations: teacher preferences are not validated by environment outcomes, and action-level scores underuse information from student rollouts, teacher rollouts, and their behavioral relationship. We therefore advocate outcome-verified teacher supervision and comparative learning over teacher-student trajectories. Based on this view, we propose Outcome-Verified Comparative Self-Distillation (OVCSD). OVCSD organizes failed student rollouts into a prefix tree, adaptively invokes a skill-conditioned teacher from student-reached states, and retains only outcome-verified successful continuations. It then applies localized comparative learning at the first state-aligned divergence and distills the post-divergence teacher suffix to transfer completion behavior. Experiments on ALFWorld and WebShop across three model scales show that OVCSD consistently outperforms skill-free RL and existing self-distillation baselines, achieving up to 29.7 and 5.4 absolute success-rate gains over the strongest baselines on ALFWorld and WebShop, respectively, while adding less than 3% privileged interaction during training.