发表机构
Simon Fraser University; NVIDIA(西蒙弗雷泽大学; 英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出扩散技能发现(DSD)方法,利用扩散模型近似熵梯度,以学习多样且可复用的运动技能,提升高维人形控制中的行为覆盖,并支持分层与零样本下游任务。
AI 中文摘要
人类通过在不同目标和情境中复用丰富的运动技能库,能够高效地学习新任务。类似的策略也可用于使模拟角色通过利用可复用的运动技能来高效执行新任务。为了支持广泛的下游任务,所学习的技能库应具有多样性,包含不同的行为以及每种行为内的空间和时间变化。学习多样技能的一种常用方法是通过最大化技能潜在变量与策略产生的状态之间的互信息。边际状态熵促进广泛的行为覆盖,而条件熵鼓励每个潜在变量产生一致的行为。然而,在高维控制问题中,直接估计边际状态熵是不可行的。因此,先前的方法依赖于间接的潜在空间近似或状态分布的粗略估计器。这些近似可能无法有效促进状态空间的广泛覆盖,导致技能的行为多样性有限,并降低对下游任务的实用性。在这项工作中,我们提出了扩散技能发现(DSD),一种通过扩散模型利用分数匹配来近似策略诱导状态分布的熵梯度的技能发现方法。由此产生的目标鼓励发现能够为高维人形控制产生更广泛行为范围的技能。所学习的技能在两个下游控制设置中被复用:通过任务特定的高层策略进行分层控制,以及通过从离线轨迹中进行潜在选择进行零样本控制。我们的实验表明,DSD比先前的技能发现方法发现了更广泛的可复用运动技能库,从而涌现出可跨下游任务复用的复杂且敏捷的行为。
英文摘要
Humans efficiently learn new tasks by reusing a rich repertoire of motor skills across different goals and contexts. A similar strategy can also be used to enable simulated characters to efficiently perform new tasks by leveraging reusable motor skills. To support a wide range of downstream tasks, the learned repertoire should be diverse, consisting of distinct behaviors as well as spatial and temporal variation within each behavior. A commonly used method for learning diverse skills is by maximizing the mutual information between skill latents and the states produced by a policy. The marginal state entropy promotes broad behavioral coverage, while the conditional entropy encourages consistent behaviors from each latent. However, directly estimating the marginal state entropy is intractable in high-dimensional control problems. Prior methods therefore rely on indirect latent-space approximations or coarse estimators of the state distribution. These approximations may not effectively promote broad coverage of the state space, resulting in skills with limited behavioral diversity and reduced utility for downstream tasks. In this work, we propose Diffusion Skill Discovery (DSD), a skill discovery method that uses a diffusion model to approximate the entropy gradient of the policy-induced state distribution through score matching. The resulting objective encourages the discovery of skills that produce a broader range of behaviors for high-dimensional humanoid control. The learned skills are reused in two downstream control settings: hierarchical control with a task-specific high-level policy and zero-shot control through latent selection from offline trajectories. Our experiments show that DSD discovers a broader repertoire of reusable motor skills than prior skill discovery methods, leading to the emergence of complex and agile behaviors that can be reused across downstream tasks.