发表机构
Drexel University(卓克索大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出基于冻结DINO编码器的潜在世界模型,通过自回归回退损失与单调代价排序损失优化,在GNM数据集上实现图像目标导航最优性能,方向误差较基线降低2.7倍,还可零样本部署于物理机器人。
AI 中文摘要
采用潜在世界模型的图像目标导航不仅需要准确的未来预测,还需要能可靠对候选动作序列排序的规划代价。我们将该代价定义为预测未来嵌入与目标嵌入之间的余弦距离,并表明糟糕的代价排序会误导交叉熵方法(Cross-Entropy Method,CEM)等基于采样的规划器。为解决该问题,我们提出一种基于冻结的DINO系列编码器构建的潜在世界模型,并用两个互补目标对其进行训练:自回归回退损失缩小训练与多步规划回退之间的差距;单调代价排序(Monotone Cost Ranking,MCR)损失直接促使受扰动程度递增的动作序列获得更高规划代价。我们还研究了基于InfoNCE的动作对比训练,发现时序排列负样本会扭曲潜在几何结构并降低规划性能。在GNM导航数据集上,我们的方法在图像目标导航性能上优于导航世界模型(Navigation World Models,NWM)、DINO-WM、OmniVLA和NoMaD,达到了最优性能,且相比使用相同编码器的DINO-WM基线,方向误差降低了2.7倍。我们还将该模型零样本部署到物理机器人上,使其在未见的室内和室外环境中遵循目标导向的路径。
英文摘要
Image-goal navigation with latent world models requires not only accurate future prediction, but also a planning cost that reliably ranks candidate action sequences. We define the cost as the cosine distance between the predicted future embedding and the goal embedding, and show that poor cost ordering can mislead sampling-based planners such as Cross-Entropy Method (CEM). To address this, we propose a latent world model built on a frozen DINO-family encoder and train it with two complementary objectives. An autoregressive rollout loss reduces the gap between training and multi-step planning rollouts, while a Monotone Cost Ranking (MCR) loss directly encourages increasingly perturbed action sequences to receive higher planning costs. We also study InfoNCE-based action-contrastive training and find that temporal permutation negatives distort the latent geometry and degrade planning performance. On the GNM navigation dataset, our method outperforms Navigation World Models (NWM), DINO-WM, OmniVLA, and NoMaD, achieving state-of-the-art image-goal navigation performance while reducing orientation error by $2.7\times$ over the same-encoder DINO WM baseline. We also deploy the model zero-shot on a physical robot, where it follows goal-directed paths in unseen indoor and outdoor environments.