DevGRU:基于碰撞感知循环模型的深度引导视觉导航
DevGRU: Depth-guided Visual Navigation using a Collision-aware Recurrent Model
浏览论文内容
中文总结 AI 辅助
针对现有视觉导航模型易在复杂室内环境碰撞的问题,提出DevGRU导航系统,结合动作与碰撞预测器提升导航性能,在9种场景实验中优于多种基线模型,且模型规模小、推理速度快
中文摘要 AI 辅助
现有视觉导航模型常致力于开发可让机器人在不同平台间通用的基础模型,但其中许多模型在复杂室内环境(尤其是结构化布局与狭窄通道)中部署时易发生碰撞。为解决该问题,我们提出基于深度图像与点目标条件的导航系统DevGRU,其采用动作预测器(AP)生成碰撞感知的未来轨迹,可有效规避近距离障碍物;结合碰撞预测器后,AP还能补偿目标位姿估计中累积的误差,主动缓解未来偏差。为评估所提方法,我们在9种不同场景中开展实验,对比了ViNT、NoMaD、NavDP这3种最先进方法,以及ViNT和NoMaD的4种额外变体。在导航性能方面,DevGRU大幅优于ViNT和NoMaD;此外,该模型可训练参数数量相对较少,在基线中推理速度最快,其模型规模比NavDP小7倍,推理时间比NavDP快17倍。
英文摘要
Existing visual navigation models often aim to develop foundation models that can generalize robot navigation across diverse platforms. However, many of these models are prone to collisions when deployed in complex indoor environments, particularly in structured layouts and narrow passages. To address this problem, we propose a depth image- and point-goal-conditioned navigation system, DevGRU. The proposed system employs an action predictor (AP) that generates collision-aware future trajectories, enabling effective avoidance of immediate obstacles. In conjunction with a collision predictor, the AP further compensates for errors accumulated in the goal pose estimation and proactively mitigates future deviations. To evaluate our method, we conducted experiments across nine different scenes and three state-of-the-art approaches - ViNT, NoMaD, and NavDP - as well as four additional variants of ViNT and NoMaD. In terms of navigation performance, DevGRU significantly outperforms ViNT and NoMaD by a large margin. In addition, the proposed model has a relatively small number of trainable parameters, resulting in the fastest inference time among the baselines, particularly outperforming NavDP by 7x in model size and 17x in inference time.