arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.25215cs.ROcs.CV

利用语义地图进行城市尺度的跨视图定位

Leveraging Semantic Maps for City-Scale Cross-View Localization

  • CSAIL, MIT(麻省理工学院计算机科学与人工智能实验室)
  • DEVCOM Army Research Laboratory (ARL)(陆军研究实验室)

机构由 AI 辅助整理,请以论文原文为准。

Ethan Fahnestock, Erick Fuentes, Philip R Osteen, Nicholas Roy

AI总结:

研究如何让机器人在未遍历环境中利用先验数据定位,提出克服提取语义信息和关联观测与先验地图两挑战的方法,即从视觉语言模型提炼轻量级匹配器,结合贝叶斯滤波器创建姿态估计,发布数据集并证明方法可泛化。

AI中文摘要:

我们希望机器人能在以前未遍历的环境中根据常见的先验数据进行定位。来自OpenStreetMap的丰富语义数据在此任务中可能有用。然而,现有方法要么忽略此语义信息,直接匹配全景图和俯视图,要么大幅压缩语义信息,只处理少量固定类别。为利用这些丰富语义信息,需克服两个挑战:一是从机器人的自我中心观测中提取有用语义信息,二是将观测信息快速与大型先验语义地图关联。我们表明视觉语言模型(VLMs)在从全景图中提取相关地标以及识别这些地标与先验俯视图地标之间的可行对应关系方面是有效的。但随着映射地标数量增加,使用VLMs提出所有对应关系的扩展性不佳。因此,我们提出从VLM中提炼出一个轻量级匹配器,它能为地图中的所有实体计算对应关系。我们用此输出形成观测似然性,并随时间与贝叶斯滤波器融合以创建姿态估计的时间序列。为支持对利用语义信息的可推广跨视图方法的进一步研究,我们发布了一个数据集,涵盖十一个环境的提取语义和评估轨迹,包括我们在波士顿暴风雪和夜间收集的全景图。我们证明,在单个城市的晴天数据上训练的我们的方法能在不同地点、光照、天气等挑战下实现泛化。代码和数据集可在指定网址获取。

英文摘要:

We want robots to localize in previously untraversed environments against commonly available prior data. Rich semantic data available from OpenStreetMap can be useful in this task. However, existing methods either ignore this semantic information, directly matching panoramas and overhead imagery, or dramatically compress the semantic information, working with a small set of fixed classes. To leverage this rich semantic information, two challenges need to be overcome. First, useful semantic information needs to be extracted from the robot's egocentric observations. Second, the observed information must be quickly associated with the large prior semantic map (e.g., up to 628 km^2). We show that VLMs are effective at both extracting relevant landmarks from panoramas, and identifying feasible correspondences between these landmarks and prior overhead landmarks. However, using VLMs to propose all correspondences scales poorly as the number of mapped landmarks increases. Instead, we propose distilling a lightweight matcher from a VLM which computes correspondences for all entities in a map. We use this output to form an observation likelihood which is fused over time with a Bayes filter to create a time series of pose estimates. To support further investigation into generalizable cross-view methods that leverage semantic information, we release a dataset of extracted semantics and evaluation trajectories spanning eleven environments, including panoramas we collected in a snowstorm and at night in Boston. We demonstrate our method, trained on a single city's fair-weather data, generalizes across location, lighting, weather, and other challenges. Code and datasets are available at https://efahnestock.github.io/loci/.

补充信息

↑