发表机构
National University of Singapore(新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出APPL方法,利用策略的结构性先验作为组合与技能间的接口,提升机器人的技能泛化与组合能力,在MetaWorld和ManiSkill任务上验证了有效性。
AI 中文摘要
从少量演示中学习的机器人通常需要两种形式的泛化。组合泛化将技能重新组合以解决新任务,而技能泛化则让每个技能背后的学习策略在新情境中发挥作用。这两者相互依赖,但在组合与它所调用的技能之间,信息会丢失。技能在何处有效,由其策略训练时所采用的结构决定,而组合仅通过一个独立的描述(如名称、指令或符号操作符)来观察技能,该描述省略了这一结构。我们的核心思想是,将每个策略的结构性先验用作组合与技能之间接口的一部分。结构性先验陈述了行为依赖于什么,例如,抓取仅依赖于夹爪相对于物体的位姿。将其内置于训练中,它塑造了策略泛化的范围;以语言表述时,它告诉组合该策略适用于何处。我们在智能体先验引导的策略学习(APPL)中实现了这一思想。一个构建智能体将完整的演示分割为可复用的技能,为每个技能提出若干结构性先验,并为每个先验训练和验证一个策略。随后,一个运行时智能体在这些特定先验的策略中进行选择,并利用它们的接口将其组合以达成新的任务目标。在MetaWorld和长时程ManiSkill任务中,APPL提升了分布外技能泛化能力,并实现了此前未见过的技能组合;消融接口信息会大幅降低性能。这些结果支持将训练时的结构性假设作为技能学习与技能组合之间桥梁的可行性。
英文摘要
Robots that learn from a few demonstrations often require two forms of generalization. Compositional generalization recombines skills to solve new tasks, and skill generalization lets the learned policy behind each skill work in new situations. The two depend on each other, yet information is lost between composition and the skills it calls. Where a skill works is determined by the structure its policy is trained with, while composition sees the skill only through a separate description, such as a name, an instruction, or a symbolic operator, that omits this structure. Our key idea is to use each policy's structural prior as part of the interface between composition and the skill. A structural prior states what a behavior depends on, for example that a grasp depends only on the gripper's pose relative to the object. Built into training, it shapes where the policy generalizes; stated in language, it tells composition where the policy applies. We instantiate this idea in Agent Priors-guided Policy Learning (APPL). A construction agent segments complete demonstrations into reusable skills, proposes several structural priors for each skill, and trains and verifies one policy per prior. A runtime agent then selects among these prior-specific policies and composes them toward new task goals using their interfaces. Across MetaWorld and long-horizon ManiSkill tasks, APPL improves out-of-distribution skill generalization and enables previously unseen skill compositions; ablating the interface information substantially reduces performance. These results support the use of training-time structural assumptions as a bridge between skill learning and skill composition.