arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

CLIP在区域地理定位任务中学习到了什么?适配后的视觉线索与场景配置探究

What Does CLIP Learn for Regional Geolocalization? Probing Visual Cues and Scene Configuration After Adaptation

Changyu Lee, Yeonsoo Park, Abdullah Alfarrarjeh, Seon Ho Kim

arXiv 2608.21761首次发表:更新:

发表机构

University of Southern California; USC IMSC; German Jordanian University; Integrated Media Systems Center (IMSC)(南加州大学; 南加州大学综合媒体系统中心; 德国约旦大学; 综合媒体系统中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对大都市区区域地理定位问题,对比CLIP不同适配方法的性能,发现编码器适配可提升邻近区域区分能力,且对场景配置的敏感性更高。

AI 中文摘要

大量街景图像集合提供了城市环境的丰富视觉信息,但从这类数据中提取细粒度地理信息仍具挑战性。特别是细粒度区域地理定位难度较大,因为邻近区域往往共享粗粒度地理线索。我们研究大都市区内的区域地理定位问题,探究预训练的CLIP特征是否足以实现区域区分,以及适配后何种视觉信息支撑性能。我们使用来自大洛杉矶地区8个区域的9085张街景图像,对比了零样本CLIP、冻结编码器读出、部分编码器更新、低秩适配(LoRA)和全微调方法。冻结读出的准确率接近零样本的39.03%,而编码器适配的准确率达到75.94%-82.10%。全微调还将到预测区域中心的平均距离从12.30公里降至3.86公里。我们通过语义线索移除、使用边缘图和模糊处理的外观缩减,以及使用补丁打乱的场景配置破坏来探究这些性能提升。适配后的模型在边缘和模糊任务上达到更高的准确率,且在补丁打乱后会改变42.92%-45.56%的预测,而冻结方法仅改变10.79%-14.60%。不过,适配并未提升外观缩减后保留的性能比例,植被和天空仍具影响力。Caltech101对照实验进一步表明,打乱敏感性并非地理定位任务所独有。总体而言,编码器适配大幅提升了邻近区域的区分能力,且与对完整场景配置的更高敏感性相关,没有证据表明仅靠粗结构就足以进行预测。这些结论涉及已知位置附近的视角变化,而非地理上不相交的泛化。

英文摘要

Large collections of street-view imagery provide rich visual information about urban environments, but extracting fine-grained geographic information from such data remains challenging. In particular, fine-grained regional geolocalization is challenging because nearby areas often share coarse geographic cues. We study regional geolocalization within a metropolitan area and ask whether pretrained CLIP features are sufficient for regional discrimination, and what visual information supports performance after adaptation. Using 9,085 street-view images from eight Greater Los Angeles regions, we compare zero-shot CLIP, frozen-encoder readouts, partial encoder updating, Low-Rank Adaptation (LoRA), and full fine-tuning. Frozen readouts remain near the 39.03% zero-shot accuracy, whereas encoder adaptation achieves 75.94-82.10%. Full fine-tuning also reduces the mean distance to the predicted region center from 12.30 km to 3.86 km. We probe these gains through semantic cue removal, appearance reduction using edge maps and blur, and scene-configuration disruption using patch scrambling. Adapted models achieve higher edge and blur accuracy and switch 42.92-45.56% of predictions after scrambling, compared with 10.79-14.60% for frozen methods. However, adaptation does not improve the fraction of performance retained after appearance reduction, while vegetation and sky remain influential. A Caltech101 control further shows that scrambling sensitivity is not unique to geolocalization. Overall, encoder adaptation substantially improves nearby-region discrimination and is associated with greater sensitivity to intact scene configuration, without evidence that coarse structure alone becomes sufficient for prediction. These conclusions concern viewpoint variation near known locations rather than geographically disjoint generalization.

Comments10 pages, 4 figures, 9 tables. Manuscript under review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑