发表机构
Zhejiang University; Shandong University(浙江大学; 山东大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对具身导航中3D场景理解与导航推理协同问题,提出SoftNav方法,通过轻量级投影仪将3D连续表示作为软令牌注入VLM隐藏空间,在HM3D-OVON上超越先前方法,且能零样本转移到其他场景,弥合表征差距实现可转移导航。
AI 中文摘要
在目标导向的具身导航中,智能体需在未知环境中定位指定目标,3D场景理解与导航推理需协同工作。当前方法通过文本向视觉语言模型(VLM)传输3D场景信息,存在表征差距。我们引入SoftNav,通过轻量级投影仪将实体级3D连续表示作为软令牌注入VLM的隐藏空间。在HM3D-OVON上,SoftNav在三个分割上的成功率分别达到74.2%/68.3%/66.7%,超越所有先前方法,且同一导航策略可零样本转移到其他场景。直接将3D场景令牌注入VLM弥合了表征差距,实现了最少训练下的可转移导航。
英文摘要
In goal-directed embodied navigation, where an agent must locate a specified target in an unseen environment, 3D scene understanding and navigation reasoning must work in concert. Current approaches transmit 3D scene information to vision-language models (VLMs) through text, suggesting a representation gap in our tested configurations; a controlled ablation confirms that direct embedding-level transfer significantly outperforms the evaluated text serialization formats. We introduce SoftNav, which injects entity-level 3D continuous representations -- one token per detected object or frontier -- into a VLM's hidden space as soft tokens through a lightweight projector. With the 3D encoder and VLM frozen, only ~1,200 samples and ~17M trainable parameters are needed. On HM3D-OVON, SoftNav achieves 74.2%/68.3%/66.7% SR across three splits, surpassing all prior methods in both SR and SPL; the same navigation policy transfers zero-shot to GOAT-Bench (67.2% SR), SG3D (47.2% s-SR), and real-world robot deployment without retraining or architectural modification. Injecting 3D scene tokens directly into VLMs bridges the representation gap, enabling transferable navigation with minimal training.
Comments8 pages, 6 figures. Accepted to IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) 2026