发表机构
KDDI Research, Inc.(KDDI研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出将CV模型与LVLM结合的混合框架,在RescueNet、FloodNet基准上评估,可准确统计灾后建筑物损毁情况,性能优于单独基准模型且仅需少量标注数据,相关资源公开。
AI 中文摘要
灾后建筑物损毁的快速准确评估至关重要,但仍是一项具有挑战性的任务。无人机(UAV)影像可提供受灾区域及时且高分辨率的视图,但现有计算机视觉(CV)模型通常需要大量标注数据集,在不同地理区域及其评估政策间的泛化能力较差,且仅适用于其训练的特定任务。大型视觉语言模型(LVLM)凭借强大的推理和泛化能力提供了有前景的替代方案,但在精确的低层次感知任务(如目标检测和精确边界框生成)上表现不足,此外,它们在特定领域任务的有效微调中通常需要大量数据。在本文中,我们提出了一种混合框架,将检测与损毁评估解耦,结合CV模型的精确性与LVLM的推理能力:CV模型首先检测建筑物并在影像上生成边界框,随后将其传递给LVLM以进行损毁分类和上下文解释。我们在两个真实基准数据集RescueNet和FloodNet上对该框架进行了评估,该框架下的最优组合可准确统计完好、部分损毁和完全损毁的建筑物,相比单独基准模型最高提升了2.1的R²值,且检测阶段仅需有限的标注数据。除报告整体性能提升外,我们还对失败场景和边缘案例进行了详细分析,为从业者提供了实用见解,并为未来工作指明了具体方向。我们的源代码和数据可通过以下存储库向研究界公开获取:this https URL
英文摘要
Rapid and accurate post-disaster building damage assessment is essential, yet remains a challenging task. Unmanned Aerial Vehicle (UAV) imagery offers a timely and high-resolution view of affected areas, but existing Computer Vision (CV) models often demand large annotated datasets, generalize poorly across geographic regions and their assessment policies, and are confined to the specific tasks they were trained for. Large Vision-Language Models (LVLMs) offer a promising alternative through their strong reasoning and generalization capabilities, but fall short on precise, low-level perception tasks such as object detection and accurate bounding box generation. Furthermore, they often require a substantial amount of data for effective fine-tuning on domain-specific tasks. In this paper, we propose a hybrid framework that decouples detection from damage assessment, combining the precision of CV models with the reasoning power of LVLMs. A CV model first detects buildings and generates bounding boxes on the image that are then passed to an LVLM for damage classification and contextual interpretation. We evaluated our framework on two real-world benchmarks: RescueNet and FloodNet. In particular, the best combination under this framework accurately counts intact, partially damaged and completely destroyed buildings, surpassing isolated baselines by up to 2.1 R^2 points, while requiring only limited annotated data for the detection stage. Beyond reporting aggregate gains, we provide a detailed analysis of failure scenarios and edge cases, offering practical insights for practitioners and concrete directions for future work. Our source code and data are publicly available to the research community via the following repository: https://github.com/ungquanghuy-kddi/VLM_GDINO.git
CommentsAccepted at ECMLPKDD 2026, 31 pages (including appendix), 18 figures