使用目标条件双模拟学习可迁移技能
Learning Transferable Skills using Goal-Conditioned Bisimulation
浏览论文内容
中文总结 AI 辅助
提出一种基于目标条件双模拟的无监督技能发现方法,学习动作感知时间表示并仅条件于关键状态特征,实现跨布局的可迁移技能,并在实证中展现强分布外泛化。
中文摘要 AI 辅助
无监督技能发现已成为利用无奖励数据集预训练通用策略的一种有前景的方法。然而,当前的技能发现方法要么需要访问专家数据,要么表现出有限的泛化能力,无法有效地迁移到以前未见过的布局中。一个关键挑战是学习能够捕捉环境时间结构同时保持对不同布局变化鲁棒的表示。为了解决这个问题,我们提出了一种学习动作感知时间表示的目标,该表示满足功能等变性属性,同时保留环境的局部时间结构。在此嵌入的基础上,我们进一步提出了使用双模拟的无监督技能发现方法,该方法通过仅将技能行为条件限制在直接影响其执行的状态特征子集上来学习可迁移技能。这强制了不同布局之间的不变行为,使技能能够有效地迁移到其他配置中。最后,通过全面的实证评估,我们展示了在给定环境中学习的技能可以有效地应用于解决各种环境布局中的下游任务,展示了强大的分布外泛化能力。
英文摘要
Unsupervised skill discovery has emerged as a promising approach for leveraging reward-free datasets to pretrain general-purpose policies. However, current skill discovery methods either require access to expert data or exhibit limited generalization, failing to transfer effectively to previously unseen layouts. A key challenge is to learn representations that capture the temporal structure of the environment while remaining robust to variations across layouts. To address this issue, we present an objective for learning action-aware temporal representations that satisfy the functional equivariance property while preserving the local temporal structure of the environment. Building upon this embedding, we further propose unsupervised skill discovery using bisimulation, which learns transferable skills by conditioning the behavior of skills exclusively on the subset of state features that directly affect their execution. This enforces invariant behavior across different layouts, enabling skills to transfer effectively to other configurations. Finally, through comprehensive empirical evaluations, we show that skills learned in a given environment can be effectively applied to solve downstream tasks in various environment layouts, demonstrating strong out-of-distribution generalization.
发表机构
- Sharif University of Technology(谢里夫理工大学)
机构由 AI 辅助整理,请以论文原文为准。