发表机构
New York University(纽约大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究揭示用视觉-语言模型测量城市变化时,同一条街道重拍会使感知分数平均变0.80分,单个样本点不可靠但数百对观测聚合后可得到可靠重建信号。
AI 中文摘要
视觉-语言模型正越来越多地被用于从重复的街景图像中测量城市变化,但其纵向可靠性尚未得到充分理解。我们测试了当街道本身未经历重大重建时,感知分数会发生多大程度的变化。我们使用来自美国五个城市435个谷歌街景视点的4648个连续时期的图像对,发现重新拍摄同一条街道会使感知分数平均变化0.80分,相当于同一城市中两条不同街道之间分数差异的66.5%。重复的模型调用几乎不会产生任何变化,而图像重新编码和提示顺序变化各占街道间差异的约五分之一。描述散射、对比度、颜色、曝光、锐度和镜面反射的六个图像统计量几乎无法解释剩余的时期间变化。仍存在约0.1分的小幅系统性漂移,且随拍摄间隔增加而增大,这与重建标签未记录的微小物理变化一致。对照实验进一步表明,当允许相机和图像属性变化时,采集条件可改变分数,且这些变化的方向取决于模型。在众包图像中,仅相机几何就会导致模型在45%的相同场景对中报告物理变化;将两张图像归一化为共同的虚拟相机可将该比例降至7.5%。尽管在单个位置层面可靠性较差,但聚合能恢复出连贯的重建信号:被判定为变化的街道更富裕、维护更好、更封闭且绿化更少。这些结果表明,视觉-语言模型对城市变化的测量在数百对观测的规模上是可靠的,但在单个样本点的规模上不可靠。
英文摘要
Vision-language models (VLMs) are increasingly used to measure urban change from repeated street-level imagery, and the results are often mapped for individual sample points. Such maps assume that a perception score changes only if the street does. We test this assumption with repeated Google Street View captures of unchanged streets in five US cities and with controlled experiments that hold the photograph, scene or camera fixed. Re-photographing an unchanged street shifts its perception score by two-thirds of the average difference between two streets in the same city. Much of this shift arises in the scoring and averages out over question orders; the rest comes from the photographs and is not predicted by image statistics of weather and light. A single image measures a place moderately well, but the change at one location barely separates redeveloped from unchanged streets. All models tested score degraded images as worse-looking streets, and in crowdsourced imagery camera differences cause false reports of physical change unless both images are rendered through a common virtual camera. After aggregation, redeveloped streets are scored as wealthier and better maintained, with more enclosure and less greenery. Streets photographed in the same capture campaign share part of their error, which averaging within a district does not remove. For urban research and planning, VLM perception scores can support comparisons between groups of streets, such as redeveloped and unchanged streets photographed in the same campaigns, but not the identification of individual streets that improved or declined.