发表机构
Indian Institute of Technology Kharagpur(印度理工学院卡哈拉格普尔分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出方向条件化策略(DCP),通过利用表示空间中的方向与距离信息改进对比强化学习,在多个导航和操作任务中提升成功率和目标接近时间,并验证其几何捕获能力。
AI 中文摘要
对比强化学习(CRL)学习估计目标可达性的表示,但其策略仍以原始目标为条件,因此并未直接利用其评论员编码的几何结构。我们引入了方向条件化策略(DCP),一种基于对CRL进行小幅修改的方法:DCP在在线训练期间选择先前访问过的状态作为路径点,并在表示空间中根据它们的方向和距离对策略进行条件化。在部署时,DCP将相同的接口直接应用于最终目标,既不需要路径点选择也不需要规划。在九个导航和操作任务中,DCP在七个任务上取得了比CRL更高的最终成功率,并在七个任务上花费更多时间接近目标。受控迷宫实验进一步表明,DCP更准确地捕获了最短路径几何结构,并且所提供的方向对行动者的行为产生了因果影响。我们识别出路径点覆盖和排序是探索的局限性,并表明学习到的候选生成在两个受控迷宫中改善了目标达成。
英文摘要
Contrastive Reinforcement Learning (CRL) learns representations that estimate goal reachability, yet its policy remains conditioned on raw goals and therefore does not directly exploit the geometry encoded by its critic. We introduce Direction-Conditioned Policies (DCP), a method built around a small modification to CRL: DCP selects previously visited states as waypoints during online training and conditions the policy on their direction and distance in representation space. At deployment, DCP applies the same interface directly to the final goal, requiring neither waypoint selection nor planning. Across nine navigation and manipulation tasks, DCP attains higher final success rates than CRL on seven tasks and spends more time near the goal on seven. Controlled maze experiments further show that DCP captures shortest-path geometry more accurately and that the supplied direction causally influences the actor's behavior. We identify waypoint coverage and ranking as limits to exploration, and show that learned candidate generation improves goal reaching in two controlled mazes.
Comments25 Pages