发表机构
University of Pennsylvania(宾夕法尼亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出VNT-PA,一种以位姿索引深度关键帧为上下文的Transformer规划器,通过位姿差异注意力实现点目标导航,在HM3D上达到93.3%成功率和90.4% SPL,优于时序基线并提升训练效率。
AI 中文摘要
学习型导航策略通常将观测作为按时间顺序排列的历史数据进行消费,通过位置编码将每个观测与其被看到的时间相关联,这使得难以重用先前环境遍历中的经验。确实重用此类经验的系统通常构建显式表示,如地图或拓扑图,并在此基础上进行规划。我们提出VNT-PA(视觉导航Transformer与位姿注意力),一种Transformer规划器,其上下文是由相机位姿索引的一组深度关键帧。以相机位姿作为位置编码,注意力依赖于关键帧之间的位姿差异而非时间顺序。VNT-PA被训练来模仿在真实场景网格上运行的最短路径规划器,通过仅使用当前位姿和目标位置查询空间上下文来预测动作。在HM3D验证场景的点目标导航中,VNT-PA达到93.3%的成功率和90.4%的按路径长度加权的成功率(SPL),在导航性能和训练效率方面均优于将相同上下文编码为时间序列或将位姿作为输入特征的基线。由于空间上下文是位姿索引的集合,测试时可以融合来自不同轨迹的帧。该规划器在定位噪声下也比在显式地图上规划的常规基线更优雅地退化。这些结果表明,带位姿标记的经验可以直接作为学习型规划器的环境表示,并且使注意力依赖于位姿差异而非时间顺序,能加速训练并改善长时程导航。
英文摘要
Learned navigation policies typically consume observations as a temporally ordered history, with positional encodings tying each observation to when it was seen, making it difficult to reuse experience from earlier traversals of an environment. Systems that do reuse such experience usually construct an explicit representation, such as a map or a topological graph, and plan on it. We propose VNT-PA (Visual Navigation Transformer with Pose Attention), a transformer planner whose context is a set of depth keyframes indexed by camera pose. With camera poses as positional encoding, attention depends on the pose differences between keyframes rather than on their temporal order. VNT-PA is trained to imitate a shortest-path planner operating on the ground-truth scene mesh, predicting actions by querying the spatial context with only its current pose and the goal position. On point-goal navigation in HM3D validation scenes, VNT-PA reaches 93.3% success and 90.4% success weighted by path length (SPL), outperforming baselines that encode the same context as a temporal sequence or treat pose as an input feature, in both navigation performance and training efficiency. Because the spatial context is a pose-indexed set, frames from different trajectories can be fused at test time. The planner also degrades more gracefully under localization noise than a conventional baseline which plans on explicit maps. These results show that pose-stamped experience can serve directly as the environment representation for a learned planner, and that making attention depend on pose differences, rather than on temporal order, speeds up training and improves long-horizon navigation.