arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.01079cs.RO

我在哪里?基于视觉-语言模型的语义地图定位用于多模态定位

Where Am I? Semantic Map Grounding via Vision-Language Models for Multi-Modal Localization

Suraj Borate, Aarav Shah, Madhu Vadali

首次发表
浏览论文内容

中文总结 AI 辅助

将机器人定位重构为语义推理任务,使用微调的Qwen2.5-VL-7B模型融合摄像头、LiDAR和语义地图预测位姿,在室内数据集上达到98.23%位置精度,并展现出跨模态互补性和对新场景的泛化能力。

中文摘要 AI 辅助

我们通过将机器人定位重构为语义推理任务而非几何估计问题,解决了GPS受限室内环境中的定位问题。受人类利用物体级线索和标注地图进行定位的启发,我们探究视觉-语言模型在给定前置摄像头图像、极坐标LiDAR扫描和俯视语义网格地图时,能否推断出机器人位姿。我们使用LoRA微调Qwen2.5-VL-7B,并附加一个轻量级回归头,直接从最终隐藏状态预测连续位姿坐标(x, y, theta),绕过文本生成。训练采用复合位置-方向损失,并在包含120,112个样本和527个场景的自定义Gazebo数据集上进行课程学习。在包含18,017个样本的分布内测试集上,模型达到98.23%的位置精度、98.00%的方向精度、96.75%的完整位姿精度,平均位置误差0.11米,平均方向误差5.7度,每样本处理时间0.62秒。在七个未见过的物体类别上,位置精度仅下降7.2个百分点,达到90.99%,支持语义空间推理而非外观记忆。在地图不完整的情况下,微调后位置精度恢复至93.72%,显示出对过时或部分地图信息的适应性。两项消融实验突出了跨模态互补性:无LiDAR时,仅使用摄像头和地图输入,位置精度保持95.06%,仅比完整系统低3.2个百分点;然而,当摄像头面对墙壁无可见物体时,LiDAR维持92.33%的位置精度,而既无LiDAR也无可见物体时精度仅为70.74%。这表明当摄像头语义信息不可用时,LiDAR成为主要定位信号,并在遮挡或稀疏布局下提供可靠的后备。

英文摘要

We address robot localization in GPS-denied indoor environments by reframing it as a semantic reasoning task rather than a geometric estimation problem. Motivated by how humans localize using object-level cues and labeled maps, we ask whether a vision-language model, given a front camera image, a polar LiDAR scan, and a top-down semantic grid map, can infer the robot pose. We fine-tune Qwen2.5-VL-7B with LoRA and attach a lightweight regression head that predicts continuous pose coordinates (x, y, theta) directly from the final hidden state, bypassing text generation. Training uses a composite position-and-direction loss with curriculum learning on a custom Gazebo dataset of 120,112 samples and 527 scenes. On the in-distribution test set of 18,017 samples, the model achieves 98.23 percent position accuracy, 98.00 percent direction accuracy, 96.75 percent full pose accuracy, a mean position error of 0.11 m, and a mean orientation error of 5.7 degrees at 0.62 s per sample. Position accuracy drops by only 7.2 percentage points on seven unseen object categories, reaching 90.99 percent, supporting semantic spatial reasoning rather than appearance memorization. With incomplete maps, fine-tuning recovers performance to 93.72 percent position accuracy, showing adaptability to stale or partial map information. Two ablations highlight cross-modal complementarity. Without LiDAR, using only camera and map inputs, position accuracy remains 95.06 percent, only 3.2 percentage points below the full system. However, when the camera sees no visible objects in a wall-facing view, LiDAR sustains 92.33 percent position accuracy, compared with 70.74 percent when neither LiDAR nor visible objects are available. This shows that LiDAR becomes the primary localization signal when camera semantics are unavailable and provides a reliable fallback under occlusion or sparse layouts.

发表机构

  • IIT Gandhinagar(印度理工学院甘地讷格尔分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑