AI 中文总结
本研究推出SIGNPOST-Bench基准,通过包含5111个反事实组的25555个图像变体,评估20个多模态大语言模型的文本-视觉冲突解决能力,发现对抗文本会显著提升定位误差,且模型对冲突文本的鲁棒性无法由干净输入的定位性能完全预测。
AI 中文摘要
多模态大语言模型(MLLMs)通过结合视觉与文本线索对真实世界场景做出有依据的预测,但现有基准很少揭示它们在这些证据源发生冲突时如何进行仲裁。我们推出SIGNPOST-Bench,这是一个用于评估文本-视觉冲突解决能力的可控反事实基准。每张源图像被转换为包含原始、空白、相似、随机和对抗变体的反事实五元组。合成的、本地化的场景文本干预被设计为保留非文本内容,从而能够成对测量定位性能的变化以及由冲突文本引入的地理目标的定向偏移。SIGNPOST-Bench包含来自四个数据集的5111个反事实组和25555个图像变体。我们评估了来自七个提供商的20个MLLM。与原始图像相比,对抗变体将中位定位误差从282公里提高到1347公里,增幅达4.8倍。在可地理编码的对抗样本中,不同模型有6.5%至20.1%的预测位于距注入目标不到50公里的范围内,且所有评估模型均表现出从空白到对抗样本的目标距离平均配对减少。兼容、不相关和冲突的文本替换对模型预测产生不同影响,而干净输入的定位性能并不能完全预测对冲突文本的鲁棒性。这些结果确立了视觉地理定位作为场景文本仲裁的连续诊断方法,并提供了一个可控框架,用于评估MLLMs如何解决冲突的多模态证据。
英文摘要
Multimodal large language models (MLLMs) make grounded predictions in real-world scenes by combining visual and textual cues, yet existing benchmarks rarely reveal how they arbitrate between these evidence sources when they conflict. We introduce SIGNPOST-Bench, a controlled counterfactual benchmark for evaluating text-vision conflict resolution. Each source image is transformed into a counterfactual quintuplet of Original, Blank, Similar, Random, and Adversarial variants. Synthetic, localized scene-text interventions are designed to preserve non-textual content, enabling paired measurements of changes in localization performance and directed shifts toward geographic targets introduced by conflicting text. SIGNPOST-Bench contains 5,111 counterfactual groups and 25,555 image variants from four datasets. We evaluate 20 MLLMs from seven providers. Compared with Original images, Adversarial variants raise median localization error from 282 km to 1,347 km, a 4.8-fold increase. Among geocodable adversarial samples, 6.5-20.1% of predictions lie less than 50 km from the injected target across models, and every evaluated model exhibits a positive mean paired reduction in target distance from Blank to Adversarial. Compatible, unrelated, and conflicting text replacements produce distinct effects on model predictions, while clean-input localization performance does not fully predict robustness to conflicting text. These results establish visual geolocation as a continuous diagnostic of scene-text arbitration and provide a controlled framework for evaluating how MLLMs resolve conflicting multimodal evidence.
Comments27 pages, 25 figures