AdaGeoVLN:面向视觉语言导航的跨表示深度与导航时间的自适应几何选择
AdaGeoVLN: Selective Geometry Across Representation Depth and Navigation Time for Vision-Language Navigation
浏览论文内容
中文总结 AI 辅助
针对视觉语言导航中几何特征利用与历史保留问题,提出AdaGeoVLN框架,通过层级GFM-VLM融合与导航感知有界记忆,在R2R-CE和RxR-CE上以单RGB流实现强性能并降低内存。
中文摘要 AI 辅助
视觉语言导航要求在随时间维持空间理解的同时,将语言与视觉观察对齐。几何基础模型(GFM)在其层级结构中暴露了中间表示,但导航策略应如何使用这些特征以及如何保留历史几何证据的问题仍未解决。我们提出\method{},一个流式视觉语言导航框架,从\textbf{表示深度}和\textbf{导航时间}两个维度解决这些问题。层级式GFM--VLM融合将GFM的早期、中期和晚期表示分别耦合到策略的连续阶段,而非反复注入最终特征。导航感知的GFM记忆根据指令相关性、几何置信度和转换新颖性,在每层有界预算下保留历史VGGT全局注意力键值状态。保留的状态在融合到策略之前为后续观察提供几何上下文。在R2R-CE和RxR-CE上,\method{}仅使用单一RGB流且无需额外的导航专用外部数据即取得了强劲性能。受控消融实验表明,在匹配的融合位置,多深度耦合显著优于重复注入最终特征。有界的导航感知保留在显著减少GFM键值内存(相对于更大内存的时间保留)的同时保持了导航性能。这些发现支持联合考察暴露给策略的几何表示以及为未来推理保留的历史证据。代码将在录用后于该https URL发布。
英文摘要
Vision-language navigation requires aligning language with visual observations while maintaining spatial understanding over time. Geometry foundation models (GFMs) expose intermediate representations throughout their hierarchy, but how navigation policies should use these features and retain historical geometric evidence remains unresolved. We introduce \method{}, a streaming VLN framework that addresses these questions across \textbf{representation depth} and \textbf{navigation time}. Hierarchical GFM--VLM fusion couples earlier, intermediate, and later GFM representations to successive policy stages instead of repeatedly injecting a terminal feature. Navigation-aware GFM memory retains historical VGGT global-attention KV states according to instruction relevance, geometric confidence, and transition novelty under a bounded per-layer budget. Retained states provide geometric context for subsequent observations before fusion with the policy. Across R2R-CE and RxR-CE, \method{} achieves strong performance using a single RGB stream without additional navigation-specific external data. Controlled ablations show that multi-depth coupling substantially outperforms repeated terminal-feature injection at matched fusion locations. Bounded navigation-aware retention preserves navigation performance while considerably reducing GFM-KV memory relative to larger-memory temporal retention. These findings support jointly examining the geometric representations exposed to the policy and the historical evidence retained for future inference. Code will be released upon acceptance at https://humanoid-research.github.io/adageovln/.
发表机构
- VinMotion, Inc., Vietnam(VinMotion公司(越南))
- University of California San Diego(加州大学圣迭戈分校)
- University of Southern California(南加州大学)
机构由 AI 辅助整理,请以论文原文为准。