arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DisasterTD:使用多模态大语言模型和跨视角地理定位的灾害地名消歧

DisasterTD: Disaster Toponym Disambiguation Using Multimodal LLMs and Cross-View Geolocalization

Wenping Yin, Ziqi Liu, Naixia Mou, Weijia Li, Danfeng Hong, Hao Li

arXiv 2607.24856首次发表:更新:

发表机构

College of Geodesy and Geomatics, Shandong University of Science and Technology; School of Environmental Science and Spatial Informatics, China University of Mining and Technology; State Key Laboratory of Information Engineering in Surveying, Mapping and Remote Sensing, Wuhan University; Tsinghua Shenzhen International Graduate School, Tsinghua University; School of Automation, Southeast University; Department of Geography, National University of Singapore(山东科技大学测绘与地理信息学院; 中国矿业大学环境与测绘学院; 武汉大学测绘遥感信息工程国家重点实验室; 清华大学深圳国际研究生院; 东南大学自动化学院; 新加坡国立大学地理系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对社交媒体图像地理参考模糊问题,提出DisasterTD框架,集成多模态大语言模型语义推理与跨视角地理定位,在飓风哈维数据集上评估,该方法优于基线,能有效进行细粒度灾害地理定位,提升不同距离下的定位准确率并减少误差。

AI 中文摘要

社交媒体图像(SMI)能提供及时且细粒度的地面视角,对态势感知和应急响应很有价值。但SMI中的地理参考常模糊不清,准确地理定位颇具挑战。为此提出DisasterTD框架,它将基于多模态大语言模型(MLLMs)的语义推理与跨视角地理定位相结合。先由MLLMs从有噪声文本输入中提取地名并生成候选地理位置,再通过SMI、遥感图像(RSI)和街景图像(SVI)的跨视角匹配来验证和细化候选结果。在飓风哈维数据集上评估,结果显示DisasterTD持续优于仅使用MLLM和仅使用跨视角的基线方法,在不同距离下均有较高定位准确率,且减少了平均和中位数误差,尤其在模糊地名上效果显著,证明了该集成方法对细粒度灾害地理定位的有效性。

英文摘要

Social media imagery (SMI) provides timely and fine-grained ground perspectives that are valuable for situational awareness and emergency response. Unlike satellite or aerial imagery, SMI can capture disaster impacts and ground-level conditions in a timely manner. However, geographic references in SMI are often vague or ambiguous, making accurate geolocalization challenging. To address this issue, we propose DisasterTD, a disaster toponym disambiguation framework that integrates multimodal large language model (MLLMs)-based semantic reasoning with cross-view geolocalization. First, MLLMs extract toponyms and generate candidate geolocations from noisy textual inputs. Then, cross-view matching between SMI, remote sensing imagery (RSI), and optionally street-view imagery (SVI) is used to verify and refine these candidate results. We evaluate DisasterTD on the Hurricane Harvey dataset, where SMI is augmented with collected RSI and SVI to construct a cross-view benchmark for disaster geolocalization. The dataset is divided into four categories based on toponym clarity and ambiguity, allowing a fine-grained performance analysis across scenarios. Results show that DisasterTD consistently outperforms MLLM-only and cross-view-only baselines without disambiguation, achieving geolocalization accuracies of 71.62% within 1000 m, 62.36% within 500 m, 57.99% within 250 m, 52.09% within 100 m, and 47.01% within 50 m, while reducing the mean and median errors to 11.33 km and 0.68 km, respectively. The largest improvements appear in ambiguous toponyms, where semantic reasoning with cross-view evidence reduces candidate dispersion and errors. These findings demonstrate the effectiveness of integrating MLLM-based candidate generation with cross-view verification for fine-grained disaster geolocalization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑