发表机构
Wuhan University; Huazhong University of Science and Technology(武汉大学; 华中科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出LG-VLN,一个仅用单目RGB的零样本视觉语言导航框架,通过共享CleanDIFT特征和LangGraph状态编排,在R2R-CE上取得21.3%成功率,验证了共享视觉表示与显式状态管理的有效性。
AI 中文摘要
连续环境下的视觉与语言导航(VLN-CE)要求在未见过的3D环境中解释自然语言指令并执行连续的底层动作。现有方法通常依赖LiDAR、全景相机或额外传感器;分离的几何建图和语义导航视觉表示可能导致长轨迹的空间-语义不一致。我们提出LG-VLN,一个具有共享视觉特征和基于LangGraph的状态编排的单目零样本框架。在线前馈3D重建网络预测深度、相机位姿和稠密点云,用于智能体姿态估计和全局地图融合。几何和导航共享稠密CleanDIFT特征:语义一致性拒绝错误的帧间对应,而目标实例约束定义视觉参考,其相似性与局部BLIP-2图像-文本相关性结合形成语义价值图。LangGraph将指令解析、几何感知、语义价值更新、路径规划、动作执行和失败恢复表示为具有条件转换、持久状态和模块化恢复机制的有向状态图。在R2R-CE验证未见分割的固定550个片段子集上,LG-VLN实现了21.3%的成功率和12.1%的按路径长度加权成功率。消融实验表明,共享语义特征改善了导航,结合视觉相似性和图像-文本相关性进一步提升了性能。结果确立了共享视觉表示和显式状态编排对于仅使用单目RGB的零样本VLN-CE的有效性。代码将公开发布以促进可复现性。
英文摘要
Continuous-environment vision-and-language navigation (VLN-CE) requires interpreting natural-language instructions in unseen 3D environments and executing continuous low-level actions. Existing methods often depend on LiDAR, panoramic cameras, or extra sensors; separate geometric-mapping and semantic-navigation visual representations can cause long-trajectory spatial-semantic inconsistencies. We propose LG-VLN, a monocular zero-shot framework with shared visual features and LangGraph-based state orchestration. An online feed-forward 3D reconstruction network predicts depth, camera poses, and dense point clouds for agent-pose estimation and global map fusion. Geometry and navigation share dense CleanDIFT features: semantic consistency rejects incorrect inter-frame correspondences, while target-instance constraints define visual references whose similarity combines with local BLIP-2 image-text relevance to form a semantic value map. LangGraph represents instruction parsing, geometric perception, semantic value updates, path planning, action execution, and failure recovery as a directed state graph with conditional transitions, persistent state, and modular recovery mechanisms. On a fixed 550-episode subset of the R2R-CE val-unseen split, LG-VLN achieves 21.3% success and 12.1% success weighted by path length. Ablations show shared semantic features improve navigation, further boosted by combining visual similarity and image-text relevance. Results establish shared visual representations and explicit state orchestration as effective for zero-shot VLN-CE using monocular RGB alone. Code will be publicly released for reproducibility.
Comments30 pages, 4 figures