发表机构
National University of Defense Technology; Shanghai Artificial Intelligence Laboratory; Shanghai Jiao Tong University(国防科技大学; 上海人工智能实验室; 上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究推出Obshazard-bench基准测试,评估多模态基础模型的实时灾害情报能力,发现现有模型在将原始多通道物理观测转化为符合决策需求的灾害推理方面存在显著局限。
AI 中文摘要
多模态大语言模型(MLLMs)正越来越多地被用于解读地球观测数据,但其支持现实世界灾害应急响应的能力仍未得到充分评估。现有遥感基准测试大多依赖静态、事后且经专家处理的产品,如网格化再分析数据,这些数据难以与灾害快速演变、需在严格时间约束下决策的实际灾害场景相匹配。为填补这一缺口,我们推出Obshazard-bench——一个用于评估MLLMs灾害情报能力的实时、观测驱动型基准测试。与以图像为中心或事后的基准测试不同,Obshazard-bench直接整合了来自不同卫星传感器的原始高频卫星探测流,以及同期地面站观测数据、历史灾害记录和社会经济指标,绕过了延迟的专家处理和物理反演流程。该基准测试覆盖60多个国家的8大类灾害及28个亚类,包含120多个历史记录的极端事件案例和数千个面向生命周期的视觉问答(VQA)样本。此外,Obshazard-bench还定义了与灾害实际工作流程一致的三阶段评估分类:用于灾前风险检测和早期预测的预测性危机预判、用于灾中灾害追踪和终止预测的主动演变推理,以及用于灾后灾害规模推算、人道主义负担估算和社会经济影响评估的多维度影响量化。对代表性通用型和地球专项基础模型的实验表明,这些模型在将原始多通道物理观测转化为具有时间依据且符合决策需求的灾害推理方面存在显著局限。
英文摘要
Multimodal Large Language Models (MLLMs) are increasingly used to interpret Earth observation data, yet their capability to support real-world disaster emergency response remains insufficiently evaluated. Existing remote sensing benchmarks largely rely on static, post-hoc, and expert-processed products, such as gridded reanalysis data, which are difficult to align with operational disaster scenarios where hazards evolve rapidly and decisions must be made under strict time constraints. To bridge this gap, we introduce Obshazard-bench, a real-time, observation-driven benchmark for evaluating disaster intelligence in MLLMs. Unlike image-centric or post-event benchmarks, Obshazard-bench directly integrates raw, high-frequency satellite sounding streams from diverse satellite sensors with concurrent ground-station observations, historical disaster records, and socio-economic indicators, bypassing delayed expert-processing and physical-inversion pipelines. The benchmark covers 8 major disaster categories and 28 sub-categories across more than 60 countries, incorporating over 120 historically documented extreme-event cases and thousands of lifecycle-oriented VQA samples. Moreover, Obshazard-bench further defines a three-stage evaluation taxonomy aligned with the operational disaster workflow: Predictive Crisis Anticipation for pre-disaster risk detection and early forecasting, Active Evolution Reasoning for in-situ disaster tracking and termination prediction, and Multi-faceted Impact Quantification for post-disaster magnitude deduction, humanitarian burden estimation, and socio-economic impact assessment. Experiments on representative general-purpose and Earth-focused foundation models reveal substantial limitations in transforming raw multi-channel physical observations into temporally grounded and decision-relevant disaster reasoning.