发表机构
Institute of Artificial Intelligence, China Telecom; Zhejiang University; Tongji University; Shanghai Jiao Tong University(中国电信人工智能研究院; 浙江大学; 同济大学; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出带3D场景条件的WNM-3D模型,通过几何编码器与适配器整合场景上下文,经多阶段训练后在GN-Bench等数据集上的闭环导航性能优于同类模型。
AI 中文摘要
近期的视觉语言导航(VLN)系统越来越多地将预训练的视觉语言模型(VLMs)适配为视觉语言动作(VLA)策略,该策略将自我中心观测和语言指令直接映射为导航动作。尽管这类策略具备语义能力,但以动作为中心的训练未明确建模智能体的视觉观测如何在其预测运动下演变。生成式世界动作模型(WAMs)可联合预测未来观测与动作,但现有用于连续VLN的WAMs未基于从观测历史推断的几何感知表示来调节联合未来视角与动作生成。本文提出WNM-3D,一种用于连续VLN的带3D场景条件的生成式世界导航模型。为将过去观测整合为持久场景上下文,一个冻结的前馈几何编码器从单目自我中心RGB历史中提取几何感知表示,一个可训练的3D场景转令牌适配器将这些表示转换为世界动作扩散Transformer令牌空间中的固定长度前缀。通过块因果注意力,该前缀调节每个未来视频动作块,为未来视角与动作生成提供共享几何上下文。我们通过在A*生成的演示上进行监督世界动作微调、在策略访问状态上进行DAgger式适配,以及基于DanceGRPO的闭环策略优化来训练WNM-3D。在GN-Bench上的实验表明,WNM-3D在闭环导航中优于强大的基于VLM的导航策略及其2D条件对应模型;在固定近目标评估集上,WNM-3D还实现了更高的流动作一致性和更低的视觉运动误差。
英文摘要
Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VLA) policies that map egocentric observations and language instructions directly to navigation actions. Although semantically capable, such action-centric training does not explicitly model how the agent's visual observations should evolve under its predicted motion. Generative world-action models (WAMs) jointly predict future observations and actions, yet existing WAMs for continuous VLN do not condition joint future-view and action generation on geometry-aware representations inferred from the observed history. We present WNM-3D, a generative World Navigation Model with 3D scene conditioning for continuous VLN. To consolidate past observations into persistent scene context, a frozen feed-forward geometry encoder extracts geometry-aware representations from the monocular egocentric RGB history, and a trainable 3D Scene-to-Token Adapter converts them into a fixed-length prefix in the token space of the world-action Diffusion Transformer. Through block-causal attention, this prefix conditions every future video-action block, providing a shared geometric context for both future-view and action generation. We train WNM-3D through supervised world-action fine-tuning on A*-generated demonstrations, DAgger-style adaptation on policy-visited states, and Counterfactual DanceGRPO refinement for closed-loop execution. Experiments on GN-Bench show that WNM-3D outperforms strong VLM-based navigation policies and its 2D-conditioned counterpart in closed-loop navigation. Stage-wise ablations further show that DAgger-SFT provides the larger success-rate gain, while Counterfactual DanceGRPO subsequently improves both navigation success and path efficiency.