AI 中文总结
针对连续环境视觉语言导航中开环执行的状态漂移问题,提出带闭环反馈的自校正世界模型SC²-WM,通过状态级计划优化与测试时模型自适应提升导航鲁棒性和泛化性。
AI 中文摘要
连续环境下的视觉语言导航(VLN-CE)要求智能体在部分可观测的条件下做出细粒度的导航决策。然而,现有多数方法依赖开环执行,缺乏在推理过程中检测并校正内部状态漂移的机制。我们提出SC²-WM,这是一种为VLN-CE引入内部反馈以实现闭环决策的自校正世界模型框架。该方法从世界模型的前瞻中推导反馈,在执行动作前进行状态级的计划优化。为应对具有挑战性的场景,我们进一步引入条件世界感知自适应,当反馈显示模型能力不足时,该机制可在测试时通过选择性更新世界模型来实现模型级校正。在标准VLN-CE基准上的实验表明,该方法提升了导航的鲁棒性与泛化性。我们的代码可在该https网址获取。
英文摘要
Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires agents to make fine-grained navigation decisions under partial observability. However, most existing methods rely on open-loop execution, lacking mechanisms to detect and correct internal state drift during inference. We propose SC$^{2}$-WM, a self-correcting world model framework that introduces internal feedback for closed-loop decision making in VLN-CE. Our method derives feedback from world-model foresight to perform state-level plan refinement before action execution. To handle challenging scenarios, we further introduce conditional world-aware adaptation, which enables model-level correction by selectively updating the world model at test time when feedback indicates model capacity insufficiency. Experiments on standard VLN-CE benchmarks demonstrate improved navigation robustness and generalization. Our code is available at https://github.com/sunrise-ikun/SC2_WM.
CommentsAccepted by ICML 2026