AUSO:从内化到利用的动作级统一技能优化
AUSO: Action-Level Unified Skill Optimization from Internalization to Utilization
浏览论文内容
中文总结 AI 辅助
AUSO是一种统一技能学习与使用的动作级优化方法,通过渐进式过程解决现有技能建模问题,在ALFWorld等数据集上提升了智能体性能与分布外泛化能力。
中文摘要 AI 辅助
技能在智能体策略演进过程中扮演不同角色:首先应提供可学习的知识,随后支持能力形成,最终仅在提升个体决策时被调用。现有方法很少对该生命周期进行建模,它们要么将技能保留在模型外部、完全内化,要么通过嘈杂的任务级成功率在内化与利用目标间进行选择。这类设计会碎片化训练,并为同一轨迹内的动作分配统一重要性,即便技能指导可能对部分决策有帮助,却会干扰其他决策。为解决这些问题,我们提出AUSO(Action-level Unified Skill Optimization,动作级统一技能优化),它通过渐进式、动作感知的优化过程统一技能学习与技能使用。训练初期,AUSO从教师指导与环境结果中联合学习,使策略在不丢失面向任务反馈的前提下获取基础技能;随后它强调基于结果的策略优化,以巩固自主解决问题的能力。随着策略成熟,AUSO在技能条件与无技能两种情境下评估每个采样动作,所得动作级信息信号与轨迹结果优势耦合,使有益的技能敏感动作获得更强更新,有害动作则被抑制。因此,技能逐渐从外部监督源过渡为决策知识,其利用方式适配于动作级收益,而强化学习仍是各阶段的共享主干。在ALFWorld、WebShop与SearchQA上的实验表明,AUSO相较于有竞争力的基线,能持续提升智能体性能与分布外泛化能力。
英文摘要
Skills play different roles as an agent's policy evolves: they should first provide learnable knowledge, then support capability formation, and finally be invoked only when they improve individual decisions. Existing methods rarely model this lifecycle. They either keep skills outside the model, fully internalize them, or select among internalization and utilization objectives through noisy task-level success rates. Such designs fragment training and assign uniform importance to actions within the same trajectory, even though skill guidance may help some decisions while distracting others. To solve these problems, we introduce AUSO (Action-level Unified Skill Optimization), which unifies skill learning and skill use through a progressive, action-aware optimization process. At the beginning of training, AUSO jointly learns from teacher guidance and environmental outcomes, enabling the policy to acquire foundational skills without losing task-oriented feedback. It subsequently emphasizes outcome-based policy optimization to consolidate autonomous problem-solving ability. As the policy matures, AUSO evaluates each sampled action under both skill-conditioned and skill-free contexts. The resulting action-level information signal is coupled with the trajectory outcome advantage, allowing beneficial skill-sensitive actions to receive stronger updates and harmful ones to be suppressed. Therefore, skills gradually transition from an external source of supervision into decision knowledge whose utilization is adapted to its action-level benefit, while reinforcement learning remains the shared backbone across all stages. Experiments on ALFWorld, WebShop, and SearchQA show that AUSO consistently improves agent performance and out-of-distribution generalization over competitive baselines.
发表机构
- University of Science and Technology of China(中国科学技术大学)
- University of New South Wales(新南威尔士大学)
- University of Chinese Academy of Sciences(中国科学院大学)
- Xi’an Jiaotong-Liverpool University(西交利物浦大学)
机构由 AI 辅助整理,请以论文原文为准。