发表机构
Nanjing University; Institute of Automation, Chinese Academy of Sciences; Nanjing University of Aeronautics and Astronautics; MAICRO(南京大学; 中国科学院自动化研究所; 南京航空航天大学; 迈克罗)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
GeoScaffold通过训练时内化几何监督,学习紧凑几何潜变量,在不增加推理开销的情况下提升连续VLN导航性能,优于纯视觉方法。
AI 中文摘要
最近的视觉-语言导航(VLN)系统越来越多地采用流式视频-LLM策略,将自我中心的RGB观测和指令直接映射到低级动作。然而,这些策略从2D预训练中继承了较弱的3D几何先验。现有的几何感知扩展在推理时持续付出代价:深度传感器、3D编码器或每步感知工具调用。我们提出GeoScaffold,一种几何监督框架,通过在训练时将几何内化到策略本身,仅支付一次代价。它首先在训练轨迹的深度图上学习一个紧凑的深度分词器并冻结它。然后,使用少量可学习的几何查询令牌微调策略,训练其隐藏状态以重建导航关键几何,如深度、连通性和可通行性。这种监督将查询状态转化为用于动作解码的紧凑几何潜变量,并通过共享权重将几何内化到骨干网络自身的表示中。如同脚手架,分词器、目标生成器和重建头在训练后被丢弃,留下骨干网络和动作接口不变。大量实验表明,GeoScaffold在连续VLN基准上持续优于领先的纯视觉导航器,为轻量级边缘部署空间感知具身导航模型提供了实用范式。
英文摘要
Recent vision-and-language navigation (VLN) systems increasingly adopt streaming Video-LLM policies that map egocentric RGB observations and instructions directly to low-level actions. Yet these policies inherit weak 3D geometric priors from 2D pretraining. Existing geometry-aware extensions charge a persistent inference-time price: depth sensors, 3D encoders, or per-step perception tool calls. We propose GeoScaffold, a geometric supervision framework that pays this price once, at training time, by internalizing geometry into the policy itself. It first learns a compact depth tokenizer on depth maps from the training trajectories and freezes it. It then fine-tunes the policy with a handful of learnable geometry query tokens, training their hidden states to reconstruct navigation-critical geometry such as depth, connectivity, and traversability. This supervision turns the query states into compact geometric latents for action decoding, and through the shared weights also internalizes geometry into the backbone's own representations. Like a scaffold, the tokenizer, target generators, and reconstruction heads are discarded after training, leaving the backbone and action interface unchanged. Extensive experiments show that GeoScaffold consistently outperforms leading vision-only navigators on continuous VLN benchmarks, offering a practical paradigm for lightweight edge deployment of spatially aware embodied navigation models.