发表机构
ETH Zurich; Google(苏黎世联邦理工学院; 谷歌公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出将建筑物损伤评估(BDA)转化为文本序列预测任务,基于通用视觉语言模型(VLM),用开源Gemma模型初步实现取得了有前景的双时相卫星图像损伤制图结果。
AI 中文摘要
传统上,建筑物损伤评估(BDA)要么采用专用网络架构,要么通过微调地理空间图像基础模型来解决。本研究探究通用视觉语言模型(VLM)能否仅通过自回归序列生成本地化建筑物并对其损伤进行分级。我们将BDA转化为预测可变长度的边界框集合,每个边界框由坐标和损伤标签指定。基于开源Gemma模型的初步实现,仅利用双时相卫星图像和合适的文本提示,就取得了有前景的损伤制图结果。
英文摘要
Conventionally, Building Damage Assessment (BDA) is tackled either with dedicated network architectures or by fine-tuning geospatial image foundation models. In this work, we ask whether a general-purpose Vision-Language Model (VLM) can localize buildings and grade their damage through autoregressive sequence generation alone. We cast BDA as predicting a variable-length set of bounding boxes, each specified by its coordinates and a damage label. Our preliminary implementation, based on the open Gemma model, achieves promising damage mapping results from only bi-temporal satellite images and a suitable text prompt.