arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12284cs.CV

GeoSEAN:东盟地区可解释的国家层面图像地理定位

GeoSEAN: Explainable Country-Level Image Geolocation for ASEAN Regions

Muhamad Syukron, Danish Rafie Ekaputra, Tintrim Dwi Ary Widhianingsih

首次发表
浏览论文内容

中文总结 AI 辅助

针对东盟国家图像地理定位难题,提出可解释流程,收集图像评估CLIP零样本分类、LightGBM和MLP分类器,MLP性能最佳,并用多种方法分析其预测以达可解释性,能准确地理定位并进行对象级视觉线索检查。

中文摘要 AI 辅助

图像地理定位旨在仅从视觉内容推断图像的地理来源。在国家具有相似城市、路边、建筑和环境特征的地区,此任务仍具有挑战性。现有模型多关注坐标级预测或分类性能,对视觉证据如何助力位置预测洞察有限。本研究为11个东盟国家提出可解释的国家层面图像地理定位流程。收集4850张图像,评估了CLIP零样本分类、LightGBM分类器和MLP分类器三种方法,MLP测试性能最佳,准确率和F1分数达85.91%。通过多种方法分析MLP分类器预测以实现可解释性,结果表明该模型能支持准确的区域图像地理定位并可对预测背后视觉线索进行对象级检查。

英文摘要

Image geolocation aims to infer the geographic origin of an image from visual content alone. However, this task remains challenging in regions where countries share similar urban, roadside, architectural, and environmental characteristics. Many existing geolocation models focus on coordinate level prediction or classification performance while providing limited insight into how visual evidence contributes to location predictions. This study presents an explainable country level image geolocation pipeline for 11 ASEAN countries. First, we collected 4,850 images from GeoGuessr style sources, Google Images, and additional street level imagery. We then evaluated three approaches on this dataset: CLIP zero shot classification, a LightGBM classifier, and an MLP classifier. The MLP achieved the best test performance, attaining an accuracy and F1 score of 85.91%. For explainability, predictions generated by the MLP classifier were analyzed post hoc using CLIP attention rollout, YOLO26 object detection on the original images, and Energy Based Pointing Game (EBPG) overlap metrics. Object level analysis indicates that frequently detected objects are not necessarily associated with the highest attention density, suggesting that object frequency and attention based visual evidence capture different aspects of a scene. These results demonstrate that the proposed model can support accurate regional image geolocation while enabling object level inspection of the visual cues underlying its predictions.

↑