arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.00031cs.CV

看见城市还是识别地点?街景图像在VLM城市感知中超越现有城市数据的价值

Seeing the City or Recognizing the Place? What Street-View Imagery Adds Beyond Existing Urban Data in VLM-Based Urban Sensing

Kaizhen Tan

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过对比街景图像与现有城市数据在七项属性上的预测表现,发现图像价值取决于视觉可读性和数据覆盖,并据此指导图像采集与地图溯源。

中文摘要 AI 辅助

街景图像越来越多地被用于推断城市属性,但仅凭预测准确性无法揭示一张照片在已有同一地点数据之外能贡献多少信息。我们将基于图像的预测与来自五个公共资源和三个VLM的七项属性的现有城市数据进行比较。同一城市单元分别使用图像、任务上下文、邻近观测和公共记录进行评估,同时通过图像替换和冲突记录来测试对数据来源的依赖。对于道路损坏、路缘坡道和房价,现有城市数据达到或超过了仅用图像的模型,而邻近的官方统计数据在人口预测上几乎与最佳图像结果相当。图像在建筑类型、建筑功能和低层楼层数方面更具信息量。对于楼层数,图像优势随距最近标注建筑距离每增加一倍而提高5.7个百分点,而对于屋顶线经常超出画面范围的高层建筑,这一优势则有所下降。模型经常遵循冲突记录。OpenFACADES楼层标注是使用OpenStreetMap楼层值生成的,其高度相关误差模式与仅用图像的重复运行结果相反。因此,街景图像的价值取决于视觉可读性和当地数据覆盖情况。将图像与现有城市数据进行比较,可以指导图像采集并阐明衍生城市地图的来源。

英文摘要

Street-view imagery is increasingly analysed with vision-language models (VLMs) to infer urban attributes, but predictive accuracy alone does not show how much a photograph contributes beyond data already available for the same place. Using three VLMs, we compare image-based predictions with existing urban data for seven attributes drawn from five public resources. Each urban unit is evaluated with imagery, with location or text context, and against non-image predictions from nearby observations or public records. We also replace images and add conflicting records to test which source the predictions follow. Nearby observations or public records matched or exceeded image-only predictions for road damage, curb ramps, population, and house price. Images were more informative for building type, building function, and the floor count of low-rise buildings. For floor count, the advantage of images over nearby OpenStreetMap labels increased by 5.7 percentage points per doubling of distance to the nearest labelled building and declined for tall buildings whose rooflines often fell outside the frame. When images and records disagreed, predictions usually moved towards the supplied record. Released OpenFACADES floor annotations, generated with OpenStreetMap floor values as input, showed the same dependence: their agreement with the reference increased with building height, whereas that of image-only reruns decreased. The value of street-view imagery therefore depends on whether an attribute is visible and how well the place is already covered by existing data. Because machine-derived labels are often reused as references, these comparisons also bear on how urban datasets are documented and evaluated.

发表机构

  • New York University(纽约大学)

机构由 AI 辅助整理,请以论文原文为准。

↑