发表机构
Inria; DI ENS, PSL; LAAS-CNRS; University of Toulouse(法国国家信息与自动化研究所; 巴黎高等师范学院计算机系,巴黎文理研究大学; 法国国家科学研究中心系统分析与体系结构实验室; 图卢兹大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出LEAP方法,通过注视不变表示和地形课程训练,仅依靠任务压力使四足机器人涌现主动感知,在危险地形导航中达到92.7%成功率,优于脚本化与被动感知。
AI 中文摘要
主动感知使自主智能体能够选择自己的视角,而非被动处理给定的视角,从而使其能够针对性地减少环境中的不确定性。学习系统通常通过手工设计的代理目标(如覆盖率或好奇心奖励)来鼓励这种行为,但这些目标可能与任务相冲突。在这项工作中,我们提出了一种无需增强任务目标即可学习涌现主动感知(LEAP)的方法。我们将问题定义为在危险地形上进行目标导向导航,且目标必须通过视觉发现。随后,我们提出了一种具有主动感知的导航策略架构,并在地形课程上对其进行训练,其中仅任务压力即可导致注视控制的涌现。这种涌现的关键在于,LEAP基于一种注视不变表示,该表示将深度图像整合到以自我为中心的信念图中。我们在留出评估场景中验证了其性能,其成功率高达92.7%,而脚本化感知的成功率为74.2%,被动感知的成功率为34.5%,并且与特权预言机的差距仅为4.6个百分点。我们验证了LEAP导航策略在无需修改的情况下,可直接应用于物理仿真中引导四足运动策略。
英文摘要
Active perception allows autonomous agents to select their viewpoints rather than passively process the viewpoints given to them, enabling them to target where to reduce uncertainty about their environment. Learned systems typically encourage this behavior with hand-designed proxy objectives, such as coverage or curiosity bonuses, that may conflict with the task. In this work, we propose a method to learn emergent active perception (LEAP) without augmentation of the task objective. We formulate the problem of goal-oriented navigation over hazardous terrains with goals that must be discovered visually. We then propose an architecture for navigation policies with active perception, and train them on a terrain curriculum where task pressure alone leads to the emergence of gaze control. Key to this emergence, LEAP works on a gaze-invariant representation that integrates depth images into egocentric belief maps. We validate its performance in held-out evaluation scenarios, where it achieves a 92.7% success rate, compared to 74.2% for scripted or 34.5% for passive perception, and comes within 4.6 points of a privileged oracle. We validate that LEAP navigation policies, unchanged, can be directly applied to steering quadrupedal locomotion policies in physics simulation.