arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UniSkill:为演化策略学习与智能体对齐的技能提案

UniSkill: Learning Actor-Aligned Skill Proposals for an Evolving Policy

Yifei Lu, Cheng Liu, Dianzhi Yu, Hui Xiang, Ji Zhang, Yuanchu Xiao, Rong Liang

arXiv 2610.10164首次发表:更新:

发表机构

Qiantang Credit; Ant Group; The Chinese University of Hong Kong(钱塘信用; 蚂蚁集团; 香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

UniSkill通过对比动作反馈实现智能体对齐的技能提案学习,无需额外回放即可优化技能库编辑,在ALFWorld和WebShop上分别取得98.4%和84.7%的成功率。

AI 中文摘要

大型语言模型智能体可以通过保留从先前交互中提炼出的可重用技能,在多个任务上实现改进。近期工作联合优化任务执行与技能提取,使策略和技能库能够共同演化。然而,随着智能体持续学习,通过后续训练步骤中的技能复用来奖励技能提案,可能会将技能收益与智能体改进相混淆,而直接测试每个提案技能则需要代价高昂的额外智能体回放。本文提出了UniSkill,它使用共享策略与环境交互,并从产生的轨迹中提出技能库编辑(添加、更新或不编辑)。具体而言,智能体从环境奖励中学习,而对比动作反馈则指导技能提案学习。该反馈通过衡量用提案技能替换检索技能如何改变当前智能体在来自同一任务的先前成功与失败轨迹之间的动作对数似然差距,提供了一种智能体对齐信号,从而避免为每个提案进行新的回放。由于当提案技能内容得分较低时,提案级反馈可能抑制原本合适的编辑操作,我们进一步应用技能编辑支持正则化以保留探索性。实验上,UniSkill取得了强劲性能,在ALFWorld上达到98.4%的成功率,在WebShop上达到84.7%,同时保持稳定的联合训练。进一步的ALFWorld实验表明,当共享策略使用较小的骨干网络时,UniSkill仍然有效。我们的实现可在以下网址获取:此https URL。

英文摘要

Large language model agents can improve across tasks by retaining reusable skills distilled from prior interactions. Recent work jointly optimizes task execution and skill extraction, enabling the policy and skillbank to co-evolve. However, as the actor continues learning, rewarding skill proposals through their reuse in subsequent training steps may conflate skill benefits with actor improvement, while directly testing each proposed skill requires costly additional actor rollouts. In this paper, we introduce UniSkill, which uses a shared policy to interact with the environment and propose skillbank edits (Add, Update, or No Edit) from the resulting trajectories. Specifically, the actor learns from environment rewards, while contrastive action feedback guides skill proposal learning. This feedback provides an actor-alignment signal by measuring how replacing the retrieved skill with a proposed skill changes the current actor's action log-likelihood gap between previously collected successful and failed trajectories from the same task, thereby avoiding new rollouts for each proposal. Since proposal-level feedback may suppress an otherwise appropriate edit operation when the proposed skill content scores poorly, we further apply skill-edit support regularization to preserve exploration. Empirically, UniSkill achieves strong performance, reaching 98.4% success on ALFWorld and 84.7% on WebShop while maintaining stable joint training. Further ALFWorld experiments show that UniSkill remains effective when the shared policy uses a smaller backbone. Our implementation is available at https://github.com/LimOkii/UniSKill.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑