发表机构
Nanjing University; Imperial College London; Sony AI; Beijing University of Posts and Telecommunications(南京大学; 帝国理工学院; 索尼人工智能公司; 北京邮电大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有生成式导航世界模型评估多候选方案成本高的问题,提出LiteNWM,通过共享视觉编码等设计提升效率,在离线评估、迁移实验及真实机器人测试中均显著优于NoMaD等方法,可用于物理机器人的未来感知规划与闭环导航。
AI 中文摘要
直接视觉导航策略能高效生成轨迹,但未显式评估其未来后果。生成式导航世界模型通过视觉回滚提供这种预见,但在评估多个候选方案时成本高昂。我们提出LiteNWM,一种潜在导航世界模型,它在候选方案间共享视觉编码,联合预测其多时间步长的动作条件未来表示,同时一个学习型评分器利用这些预测选择轨迹。在RECON、SCAND和SACSoN的离线评估中,LiteNWM相对于NoMaD+NWM-XL将宏观平均轨迹误差降低17.56%,并在RTX 5090上实现128.00倍的端到端加速。同一评分器从NoMaD迁移到MBRA时无需针对提议者重新训练,将MBRA的宏观平均轨迹误差降低16.2%。在未见过的室内外环境的真实机器人实验中,LiteNWM相对于NoMaD将导航成功率从43.3%提升至83.3%。这些结果表明,LiteNWM可部署在物理机器人上实现具未来感知的规划与闭环导航。
英文摘要
Direct visual navigation policies generate trajectories efficiently but do not explicitly evaluate their future consequences. Generative navigation world models provide this foresight through visual rollouts, which are costly when evaluating multiple candidates. We present LiteNWM, a latent navigation world model that shares visual encoding across candidates and jointly predicts their action-conditioned future representations at multiple horizons, while a learned scorer uses these predictions to select trajectories. In offline evaluations on RECON, SCAND, and SACSoN, LiteNWM reduces macro-averaged trajectory error by 17.56% relative to NoMaD+NWM-XL and achieves a 128.00-fold end-to-end speedup on an RTX 5090. The same evaluator transfers from NoMaD to MBRA without proposer-specific retraining, reducing MBRA's macro-averaged trajectory error by 16.2%. In real-robot experiments in unseen indoor and outdoor environments, LiteNWM improves navigation success from 43.3% to 83.3% relative to NoMaD. These results demonstrate that LiteNWM can be deployed for future-aware planning and closed-loop navigation on a physical robot.
Comments12 pages, 11 figures, including appendix