发表机构
Georgia Institute of Technology; Georgia Tech Research Institute(佐治亚理工学院; 佐治亚理工研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出仅以奖励和动作为条件的策略,实现观测空间完全不同的环境间零样本迁移,并验证其在多种环境中的有效性及对观测策略训练的指导作用。
AI 中文摘要
本文探讨了一种基于奖励的策略,以实现源环境和目标环境之间观测空间完全不同的零样本迁移。虽然人类能展现出令人印象深刻的适应能力,但深度神经网络策略往往难以适应新环境,并且需要大量样本才能成功迁移。相反,我们提出了一种仅以奖励和动作作为条件的新型基于奖励的策略,从而能够对观测完全不同的新环境进行零样本适应。我们讨论了基于奖励的策略所面临的挑战和可行性,然后提出了一种实用的训练算法。我们证明了奖励策略可以在三种不同环境(Pointmass、Cartpole 和 2D Car Racing)中训练,并零样本迁移到完全不同的观测,例如不同的调色板或 3D 渲染,或 Habitat-Sim 中的 Stretch 机器人导航。我们还证明了基于奖励的策略可以进一步指导目标环境中基于观测的策略的训练。
英文摘要
This paper explores a reward-based policy to achieve zero-shot transfer between source and target environments with completely different observation spaces. While humans can demonstrate impressive adaptation capabilities, deep neural network policies often struggle to adapt to a new environment and require a considerable amount of samples for successful transfer. Instead, we propose a novel reward-based policy only conditioned on rewards and actions, enabling zero-shot adaptation to new environments with completely different observations. We discuss the challenges and feasibility of a reward-based policy and then propose a practical algorithm for training. We demonstrate that a reward policy can be trained within three different environments, Pointmass, Cartpole, and 2D Car Racing, and transferred to completely different observations, such as different color palettes or 3D rendering, or Stretch robot navigation in Habitat-Sim, in a zero-shot manner. We also demonstrate that a reward-based policy can further guide the training of an observation-based policy in the target environment.
CommentsWebsite: https://morganbyrd03.github.io/reward_based_policies/