arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

解码灾难:基于视觉语言模型与众包影像的多任务地理空间推理用于灾害制图

Decoding the Disaster: Multi-Task Geospatial Reasoning with Vision-Language Models and Crowdsourced Imagery for Disaster Mapping

Wenping Yin, Fabian Desuer, Ziqi Liu, Naixia Mou, Weijia Li, Pedram Ghamisi, Xiao Xiang Zhu, Hao Li

arXiv 2610.00302首次发表:更新:

发表机构

Shandong University of Science and Technology; Technical University of Munich; Wuhan University; Tsinghua University; Helmholtz-Zentrum Dresden-Rossendorf; University of Iceland(山东科技大学; 慕尼黑工业大学; 武汉大学; 清华大学; 亥姆霍兹德累斯顿罗森多夫研究中心; 冰岛大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对众包灾害影像地理定位难、解释难的问题,提出多任务框架GRDisaster,利用视觉语言模型结合跨视角融合与空间推理指标,实现灾害制图与损害评估,提升可解释性。

AI 中文摘要

众包影像为灾害制图提供了及时、细粒度、街道层面的观测,在应急响应期间补充了传统的遥感影像(RSI)。然而,此类影像通常是非结构化的、空间上模糊的,且缺乏可靠的地理元数据,使得人工地理定位和解释劳动密集且难以扩展。本工作提出了一种多任务地理空间推理灾害制图框架,即GRDisaster,以考察视觉语言模型(VLMs)在理解、地理定位和推理众包灾害影像方面的潜力。GRDisaster基于一个新整理的基准数据集构建,该数据集源自PhotoMappers,包含26,340张图像,组织为经过人工验证的自愿地理信息(VGI)、街景影像(SVI)、RSI跨视角三元组,覆盖2018年至2024年的多个灾害事件。该框架将确定性和概率性跨视角地理定位与多视角融合相结合,以将VGI图像与地理参考的SVI和RSI关联起来。它引入了两组空间推理指标,用于跨视角地理定位验证和灾害损害评估。这些指标利用结构、环境和全局场景线索来验证跨视角对应关系,并利用专家验证的注释来评估灾害严重程度,从而提高了VLM输出的可解释性。据我们所知,本研究首次系统性地调查并提供了统一评估框架,以检验基于VLM的空间推理如何通过跨视角地理定位验证、可解释的空间推理和灾害感知的严重性评估,将众包灾害影像转化为可操作的地理空间人工智能(GeoAI)。

英文摘要

Crowdsourced imagery provides timely, fine-grained, street-level observations for disaster mapping, complementing conventional remote sensing imagery (RSI) during emergency response. However, such imagery is often unstructured, spatially ambiguous, and lacks reliable geographic metadata, making manual geolocalization and interpretation labor-intensive and difficult to scale. This work proposes a multi-task Geospatial Reasoning Disaster mapping framework, namely GRDisaster, to examine the potential of vision-language models (VLMs) in understanding, geolocalizing, and reasoning over crowdsourced disaster imagery. GRDisaster is built on a newly curated benchmark dataset derived from PhotoMappers, comprising 26,340 images organized into human-validated volunteered geographic information (VGI), street-view imagery (SVI), RSI cross-view triplets covering multiple disaster events from 2018 to 2024. The framework combines deterministic and probabilistic cross-view geolocalization with multi-view fusion to associate VGI images with georeferenced SVI and RSI. It introduces two sets of spatial reasoning indicators for cross-view geolocalization validation and disaster damage assessment. These indicators use structural, environmental, and global-scene cues to validate cross-view correspondences and visually observable damage evidence with expert-verified annotations to assess disaster severity, improving the interpretability of VLM outputs. To our knowledge, this study provides the first systematic investigation and unified evaluation framework for examining how VLM-based spatial reasoning can transform crowdsourced disaster imagery into actionable geospatial artificial intelligence (GeoAI) through cross-view geolocalization validation, interpretable spatial reasoning, and damage-aware severity assessment.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑