GeoArbiter:面向遥感多模态大语言模型的可验证性导向 grounding 方法
GeoArbiter: Verifiability-Guided Grounding for Remote-Sensing Multimodal LLMs
浏览论文内容
中文总结 AI 辅助
该研究针对遥感多模态大语言模型的事实断言问题,提出无需训练的 GeoArbiter 流程,通过仅注入图像无法验证的地理事实,有效降低了模型的断言级幻觉并提升了对冲突记录的鲁棒性。
中文摘要 AI 辅助
遥感多模态大语言模型(MLLMs)常断言图像无法证实的事实,例如设施的身份或功能。基于坐标的地理检索可补充这类缺失知识,在三个开源 MLLMs 上提升 fMoW 土地利用准确率 12.06 至 17.19 个百分点。但检索记录也可能与可见证据矛盾,研究发现模型即便图像具备决定性证据时仍常采信记录。因此提出源可信度应取决于跨模态可验证性:地理记录对图像无法验证的属性最有用,对与视觉可验证属性相悖的情况最具风险。引入 GeoArbiter,一种无需训练的流程,通过仅注入图像无法验证的地理事实来实现该原则。与跨属性类型泄露并偏置是/否响应的仲裁提示不同,内容级过滤保留了全检索准确率增益的 84.69% 至 87.15%,在盲源评判下将断言级幻觉降低 9.58% 至 26.34%,并提升了三个模型对冲突记录的鲁棒性。这些结果表明,可验证性导向的内容选择是将遥感 MLLMs 建立在易出错地理知识上的简单有效机制。
英文摘要
Remote-sensing multimodal large language models (MLLMs) often assert facts that imagery cannot establish, such as a facility's identity or function. Coordinate-keyed geographic retrieval can supply this missing knowledge, improving fMoW land-use accuracy by 12.06--17.19 points across three open MLLMs. However, retrieved records can also contradict visible evidence, and we find that models frequently follow the records even when the image is decisive. We argue that source trust should therefore depend on \emph{cross-modal verifiability}: geographic records are most useful for attributes the image cannot verify and most dangerous when they dispute visually verifiable attributes. We introduce GeoArbiter, a training-free pipeline that operationalizes this principle by injecting only image-unverifiable geographic facts. Unlike arbitration prompts, which leak across attribute types and bias yes/no responses, content-level filtering preserves 84.69--87.15\% of the full-retrieval accuracy gain, reduces claim-level hallucination by 9.58--26.34\% under a source-blinded judge, and improves robustness to conflicting records across all three models. These results identify verifiability-guided content selection as a simple, effective mechanism for grounding remote-sensing MLLMs in fallible geographic knowledge.
发表机构
- University of Minnesota, Twin Cities(明尼苏达大学双城分校)
机构由 AI 辅助整理,请以论文原文为准。