发表机构
Georgia Institute of Technology; Simon Fraser University; NVIDIA(佐治亚理工学院; 西蒙弗雷泽大学; 英伟达)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出GGSD框架,通过游戏自我对弈发现人类可玩的离散运动技能,使人类能组合技能解决未见任务,无需额外训练。
AI 中文摘要
我们提出了游戏引导的技能发现(GGSD)框架,该框架利用游戏中的自我对弈来发现人类可直接操控的运动技能。可玩技能提供了一种紧凑的抽象,通过少量学习到的行为而非底层动作来控制具身智能体。为了有效,这些技能应在语义上区分、可解释且具有表现力;而现有的无监督技能发现方法往往无法同时满足这些属性。GGSD通过将技能发现植根于竞争性游戏来实现这些目标。一个分层智能体与自身的过往版本竞争,高层策略从一个小型离散技能集中选择,而技能条件化的低层策略学习相应的行为。训练后,人类可以替代高层策略,直接通过相同的离散技能控制智能体。尽管高层动作数量少,技能转换产生了涌现的组合行为,将表现力扩展到单个原始技能之外。在Ant、Franka机械臂和Unitree G1环境中,我们展示了GGSD产生人类可玩的技能,人类可以组合这些技能来解决未见过的任务,如迷宫和立方体推动,而无需额外训练。交互式演示可在以下网址获取:此https URL。
英文摘要
We present Game-Guided Skill Discovery (GGSD), a framework that uses self-play in games to discover motor skills that are directly playable by humans. Playable skills provide a compact abstraction for controlling embodied agents through a small set of learned behaviors rather than low-level actions. To be effective, these skills should be semantically distinct, interpretable, and expressive; properties that existing unsupervised skill-discovery methods often fail to achieve simultaneously. GGSD achieves these desiderata by grounding skill discovery in competitive gameplay. A hierarchical agent competes against its past selves, with a high-level policy selecting from a small discrete skill set and a skill-conditioned low-level policy learning the corresponding behaviors. After training, a human can replace the high-level policy and directly control the agent through the same discrete skills. Despite the small number of high-level actions, skill transitions give rise to emergent combo behaviors, expanding expressivity beyond individual primitives. Across Ant, Franka-arm, and Unitree G1 environments, we show that GGSD produces human-playable skills that humans can compose to solve unseen tasks, such as Maze and CubePush, without additional training. An interactive demo is available at https://ggsd-demo.github.io.