GeoContext:视觉语言地理定位中的一个上下文阶梯、两种失败模式:对用户提供位置上下文的平面依赖与对位置断言的虚假确认
GeoContext: One Context Ladder, Two Failure Modes in Vision-Language Geolocation: Flat Reliance on User-Provided Location Context and False Confirmation of Location Claims
AI总结:
提出GeoContext基准,包含GeoHint和GeoVerify任务,通过上下文阶梯评估五个视觉语言模型,发现模型对用户位置提示的平面依赖和虚假确认两种失败模式,并发布相关资源。
AI中文摘要:
视觉地理定位基准通常要求模型判断图像拍摄地点,而未考虑用户经常提供的位置上下文。我们引入GeoContext,一个支持两个互补任务的资源:GeoHint,在给定真实但粗略的位置提示下进行开放式定位;以及GeoVerify,对图像是否在声称地点150米范围内拍摄进行二元验证。GeoContext通过根据距离和可参考性对附近参考点进行分层来构建上下文阶梯,使图像保持固定而提供的上下文变化。该基准覆盖30个城市的109个地点,并使用21,933个GeoHint响应和6,270个GeoVerify响应评估了五个视觉语言模型。我们的评估揭示了三个主要模式。首先,提示重复率在不同可参考性层级间仅变化1.5个百分点,在不同距离带间变化小于3个百分点,而由此产生的定位误差随提示距离稳步增加。其次,行为强烈依赖于无上下文性能:在无上下文准确率低的地点,定位误差与提示距离的中位数比率约为1.00,而在准确率较高的地点,该比率在0.24至0.69之间。在纠正了地点分组程序引入的偏差后,五个模型中只有一个在给定附近提示时保留负的准确率估计。第三,在GeoVerify中,对于刚超过150米容差的诱饵,没有模型达到d'=1。当灵敏度与响应偏差分离时,模型排名也会改变,并且83.8%的虚假接受以至少0.8的置信度报告。我们发布了该基准、构建流程、审计决策和评分代码。
英文摘要:
Visual geolocation benchmarks typically ask a model where an image was captured without accounting for the location context that users often provide. We introduce GeoContext, a resource supporting two complementary tasks: GeoHint, open-ended localization given a true but coarse location hint, and GeoVerify, binary verification of whether an image was taken within 150 m of a claimed place. GeoContext constructs a context ladder by stratifying nearby reference points according to distance and referenceability, allowing the image to remain fixed while the supplied context varies. The benchmark covers 109 sites in 30 cities and evaluates five vision-language models using 21,933 GeoHint responses and 6,270 GeoVerify responses. Our evaluation reveals three main patterns. First, hint repetition varies by only 1.5 percentage points across referenceability tiers and by less than 3 points across distance bands, while the resulting localization error increases steadily with hint distance. Second, behavior depends strongly on no-context performance: at sites with low no-context accuracy, the median ratio between localization error and hint distance is approximately 1.00, whereas at higher-accuracy sites it ranges from 0.24 to 0.69. After correcting for bias introduced by the site grouping procedure, only one of the five models retains a negative accuracy estimate when given a nearby hint. Third, in GeoVerify, no model reaches d' = 1 for decoys immediately beyond the 150 m tolerance. Model rankings also change when sensitivity is separated from response bias, and 83.8% of false acceptances are reported with confidence of at least 0.8. We release the benchmark, construction pipeline, audit decisions, and scoring code.