基于偏好优化与可解释视觉-语言推理的可靠多模态灾害严重程度评估研究
Towards reliable multimodal disaster severity assessment through preference optimization and explainable vision-language reasoning
浏览论文内容
中文总结 AI 辅助
本研究提出整合SFT与DPO的两阶段多模态训练框架,经跨模型验证可提升灾害评估的准确率、解释质量及对轻度灾害的检测,为应急管理提供可靠多模态系统的可复现路径。
中文摘要 AI 辅助
可靠的灾害损失评估需要模型既能提供准确预测,又能给出透明解释。然而,现有多模态方法受限于标注数据稀缺,且对推理质量的评估不足。本研究提出一个两阶段训练框架,在统一的数据构建流程中整合监督微调(SFT)与直接偏好优化(DPO)。通过单一的人在回路(HITL)标注工作流,衍生出两个互补数据集:ReasoningSet,包含用于SFT的经验证的推理依据;PreferenceSet,包含用于基于DPO的对齐的配对推理依据。该框架采用自动指标、基于模型的评分及人工排序,同时评估分类性能与解释质量。实验结果显示,与基线相比,SFT将准确率从73.64%提升至78.29%,宏F1提升29%,解释质量提升约25%;后续DPO对齐进一步提升了PreferenceSet上的可解释性。在InternVL-3-8B与LLaVA-1.5-7B上的跨模型验证,证明了该方法的稳健性与可推广性。所提框架提升了对代表性不足的轻度灾害案例的检测,减少了高风险误分类,强化了模型推理与人类判断的对齐,为应急管理开发可靠的多模态系统提供了可复现的路径,这类系统能提供可审计、可行动的灾害洞察。
英文摘要
Reliable disaster damage assessment requires models that provide both accurate predictions and transparent explanations. However, existing multimodal approaches are limited by scarce annotated data and insufficient evaluation of reasoning quality. This study proposes a two-stage training framework that integrates Supervised Fine-Tuning (SFT) and Direct Preference Optimization (DPO) within a unified data construction pipeline. From a single Human-in-the-Loop (HITL) annotation workflow, two complementary datasets are derived, namely ReasoningSet, which contains validated rationales for SFT, and PreferenceSet, which comprises paired rationales for DPO-based alignment. The framework evaluates both classification performance and explanation quality using automatic metrics, model-based scoring, and human ranking. Experimental results show that SFT improves accuracy from 73.64% to 78.29% and increases Macro-F1 by 29% compared to the baseline, while explanation quality improves by approximately 25%. Subsequent DPO alignment further enhances interpretability on the PreferenceSet. Cross-model validation on InternVL-3-8B and LLaVA-1.5-7B demonstrates the robustness and generalizability of the approach. The proposed framework improves detection of underrepresented mild damage cases, reduces high-risk misclassifications, and strengthens alignment between model reasoning and human judgment. Overall, it provides a reproducible pathway to develop reliable multimodal systems that deliver auditable, actionable disaster insights for emergency management.
发表机构
- University of Oulu(奥卢大学)
- Institute of Engineering and Management(工程与管理学院)
机构由 AI 辅助整理,请以论文原文为准。