arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DM-KG:一种提升街景图像中视觉语言模型空间认知的新方法

DM-KG: A Novel Method for Boosting Spatial Cognition of Vision-Language Models in Street View Imagery

Xinyue Xu, Zheng Zhang, Kunyang Ma, Ge Zhu, Lianshuai Cao, Lei Wang, Zixuan Li, Yi Cheng

arXiv 2607.12319首次发表:更新:

发表机构

Institute of Surveying and Mapping, Information Engineering University; Institute of Geographic Sciences and Natural Resources Research, Chinese Academy of Sciences(信息工程大学测绘学院; 中国科学院地理科学与资源研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对视觉语言模型在街景图像空间认知方面的不足,提出DM-KG框架,通过提取实体关系、结合全景分割与深度估计计算坐标并编码知识图,有效提升模型空间推理准确性,降低误差,为地理视觉问答提供新框架。

AI 中文摘要

随着视觉语言模型(VLMs)越来越多地应用于地理空间问答和视觉场景理解,提高其在街景图像上进行复杂逻辑推理的空间认知能力成为关键研究重点。然而,现有VLMs在感知现实世界街景场景中的物体位置、距离和方向时经常出现“空间语义幻觉”,且这些错误难以追踪和校准,阻碍其在地理空间任务中的实际部署。为应对这一挑战,本研究提出DM-KG(方向-度量知识图),一个基于结构的街景图像空间表示框架。通过从单个二维图像中明确提取实体间的方向和度量关系,该框架借助结构化知识图提高VLMs的空间推理准确性。具体而言,将全景分割与度量深度估计相结合以稳健计算实体级三维空间坐标,然后将实体对的时钟方位和欧几里得距离编码到JSON格式的知识图中,注入VLM作为显式几何先验来指导空间推理。在公共空间问答基准上的实验结果表明,DM-KG将距离估计中的平均绝对误差降低了31.1%,方向判断中的平均角度误差降低了65.8%,同时保持了较高的问答成功率。本研究通过建立完整的增强推理管道,显著提高了VLMs在街景场景中的空间认知能力,为开放环境中的地理视觉问答(GeoVQA)提供了一个灵活、通用且可解释的框架。

英文摘要

As vision-language models (VLMs) are increasingly deployed in geospatial question answering and visual scene understanding, improving their spatial cognition capability on street view imagery for complex logical reasoning has emerged as a key research priority. However, existing VLMs frequently suffer from "spatial semantic hallucinations" when perceiving object locations, distances, and directions in real-world street view scenes. Furthermore, such errors are often recalcitrant to tracing and calibration, posing a critical bottleneck for their practical deployment in geospatial tasks. To address this pressing challenge, this study proposes DM-KG (Direction-Metric Knowledge Graph), a structurally grounded spatial representation framework for street view imagery. By explicitly extracting directional and metric relationships between entities from a single 2D image, this framework enhances the spatial reasoning accuracy of VLMs through a structured knowledge graph. Specifically, we integrate panoptic segmentation with metric depth estimation to robustly compute entity-level 3D spatial coordinates. Subsequently, we encode the clock azimuths and Euclidean distances of entity pairs into a JSON-formatted knowledge graph, which is injected into the VLM as an explicit geometric prior to guide spatial reasoning. Experimental results on public spatial question-answering (QA) benchmarks demonstrate that DM-KG reduces the mean absolute error (MAE) in distance estimation by 31.1% and the mean angular error in direction judgment by 65.8%, while simultaneously maintaining a high QA success rate. By establishing a complete, augmented reasoning pipeline, this research significantly improves the spatial cognitive capabilities of VLMs in street view scenarios, thereby providing a flexible, generalized, and interpretable framework for geographic visual question answering (GeoVQA) in open environments.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑