InsightMap:面向具身多模态推理的结构化空间建模
InsightMap: Structured Spatial Modeling for Embodied Multimodal Reasoning
浏览论文内容
中文总结 AI 辅助
InsightMap通过俯视地图作为空间记忆和预测目标,联合学习导航与地图生成,在R2R-CE和RxR-CE上分别取得56.9%和54.9%的成功率,并提升静态空间任务性能。
中文摘要 AI 辅助
语言引导的导航需要将部分观测与持久的空间参考联系起来,并学习动作如何改变该表示。我们提出了InsightMap,一个使用俯视地图作为显式空间记忆和动作条件预测目标的框架。历史视图被链接到标记的地图位置,一个共享的多模态主干联合学习导航动作预测和动作后地图生成。地图预测提供辅助训练监督,而导航推理从观测到的空间上下文解码动作。一个对齐的RGB-D数据管道支持导航、视觉问答、情境推理和3D定位的通用接口。在R2R-CE和RxR-CE的验证-未见分割上,InsightMap分别实现了56.9%和54.9%的成功率(SR)。添加地图预测监督使R2R-CE的SR提高了4.3个百分点,按路径长度加权的成功率(SPL)提高了3.2个百分点。在静态空间任务上,InsightMap在ScanQA上达到103.7 CIDEr,在SQA3D上达到60.1%的精确匹配准确率,在ScanRefer上使用检测到的物体提议在0.5 IoU下达到53.1%的定位准确率。在Unitree Go2上,它在走廊、实验室和办公室环境中优于NaVid和NaVILA。
英文摘要
Language-guided navigation requires connecting partial observations to a persistent spatial reference and learning how actions change that representation. We introduce InsightMap, a framework that uses top-down maps as both explicit spatial memory and action-conditioned prediction targets. Historical views are linked to labeled map locations, and a shared multimodal backbone jointly learns navigation action prediction and post-action map generation. Map prediction provides auxiliary training supervision, while navigation inference decodes actions from the observed spatial context. An aligned RGB-D data pipeline supports a common interface for navigation, visual question answering, situated reasoning, and 3D grounding. On the validation-unseen splits of R2R-CE and RxR-CE, InsightMap achieves success rates (SR) of 56.9% and 54.9%, respectively. Adding map-prediction supervision improves R2R-CE SR by 4.3 and success weighted by path length (SPL) by 3.2 percentage points. On static spatial tasks, InsightMap achieves 103.7 CIDEr on ScanQA, 60.1% exact-match accuracy on SQA3D, and 53.1% grounding accuracy at 0.5 IoU on ScanRefer with detected object proposals. On Unitree Go2, it outperforms NaVid and NaVILA in hallway, lab, and office environments.
发表机构
- University of Manchester(曼彻斯特大学)
机构由 AI 辅助整理,请以论文原文为准。