CanonNav:跨平台视觉导航中解耦导航行为与相机几何
CanonNav: Disentangling Navigation Behavior from Camera Geometry in Cross-Platform Visual Navigation
- Seoul National University(首尔大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
CanonNav框架解耦导航行为与相机几何,引入相机几何规范化并结合规划监督,仅用RGB推理便在多场景视觉导航中优于相关基线与RGB-D方法。
AI中文摘要:
尽管视觉导航通过跨平台演示的模仿学习取得了进展,但充分利用此类数据仍具挑战。首先,直接从图像-轨迹对学习会将导航行为与依赖平台的相机几何纠缠在一起,这阻碍了一致学习,因为策略需从视觉观察中隐式推断相机几何,这本质上是不适定问题。其次,从演示轨迹进行模仿学习会捕获专家选定的运动,但该运动背后的中间决策是隐式的。为解决这些问题,我们提出CanonNav,这是一种视觉导航框架,可将导航行为与相机几何解耦,并将互补规划监督纳入跨平台演示的学习中。CanonNav引入相机几何规范化,将视觉观察和轨迹转换为相机一致的表示空间。基于此表示,我们使用离线可通行性估计器的伪标签推导安全性和局部进度监督:安全性监督惩罚不安全轨迹,局部进度监督引导机器人应前进的位置。在不同相机配置和环境下的实验表明,尽管推理时仅使用RGB,CanonNav仍始终优于基于RGB的基线,甚至在具有挑战性的场景中超过基于RGB-D的方法。
英文摘要:
While visual navigation has advanced through imitation learning from cross-platform demonstrations, fully leveraging such data remains challenging. First, directly learning from image-trajectory pairs entangles navigation behavior with platform-dependent camera geometry. This hinders consistent learning by forcing the policy to implicitly infer camera geometry from visual observations, an inherently ill-posed problem. Second, imitation learning from demonstrated trajectories captures the expert's chosen motion but leaves the intermediate decisions underlying that motion implicit. To address these issues, we propose CanonNav, a visual navigation framework that disentangles navigation behavior from camera geometry and incorporates complementary planning supervision into learning from cross-platform demonstrations. CanonNav introduces camera geometry canonicalization, which transforms visual observations and trajectories into a camera-consistent representation space. Building on this representation, we derive safety and local-progress supervision using pseudo-labels from an offline traversability estimator. Safety supervision penalizes unsafe trajectories, while local-progress supervision guides where the robot should advance. Experiments across diverse camera configurations and environments show that, despite using only RGB at inference, CanonNav consistently outperforms RGB-based baselines and even surpasses RGB-D-based methods in challenging scenarios.