GeoNLI - 卫星影像的自然语言解释器
GeoNLI - A Natural Language Interpreter for Satellite Imagery
- Indian Institute of Technology, Bombay(印度理工学院孟买分校)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对遥感图像多任务理解,提出统一模块化流程,集成SAM变体与多模态LLM,通过多数投票在描述、VQA和定位上分别取得82%、83.32%和64.94%的准确率。
AI中文摘要:
多模态多任务模型在遥感数据集上表现出强大的性能。然而,由于这些模型在异构数据上训练且在不同任务间存在差异,设计一个在图像描述(captioning)、视觉问答(VQA)和视觉定位(visual grounding)方面均表现良好的统一模型仍然具有挑战性。在这项工作中,我们在VRS Bench和NWPU-VHR-10数据集上评估了多个模型。EarthMind模型在图像描述和视觉问答方面均展现出强劲结果。对于视觉定位,我们提出了多种流程——RemoteSAM-SAM-v1、RemoteSAM-SAM-v2和DiffuSAM——并最终采用跨EarthMind、RemoteSAM、SAM3、Falcon、RemoteSAM-SAM3-v1、RemoteSAM-SAM3-v2和DiffuSAM预测的多数投票集成方法。我们的统一模块化流程将先进的SAM变体与多模态LLM相结合,以联合执行图像描述、视觉问答和视觉定位。它在图像描述上达到82%的准确率,在视觉问答上达到83.32%,其中二元、数值和语义问题类型分别达到90.94%、52.04%和92.06%的准确率。对于视觉定位,它达到64.94%的准确率。通过将多样化的VLM与我们定制的RemoteSAM-SAM3模型通过集成多数投票相结合,该系统相比任务特定方法提供了更准确且一致的遥感理解。
英文摘要:
Multi-modal multitasking models have shown strong performance on remote sensing datasets. However, because these models are trained on heterogeneous data and vary across tasks, designing a unified model that performs well in captioning, visual question answering (VQA), and visual grounding remains challenging. In this work, we evaluate several models on the VRS Bench and NWPU-VHR-10 datasets. The EarthMind model demonstrates strong results in both captioning and VQA. For grounding, we propose multiple pipelines - RemoteSAM-SAM-v1, RemoteSAM-SAM-v2, and DiffuSAM - and ultimately adopt a majority-voting ensemble across EarthMind, RemoteSAM, SAM3, Falcon, RemoteSAM-SAM3-v1, RemoteSAM-SAM3-v2, and DiffuSAM predictions. Our unified, modular pipeline integrates advanced SAM variants with multimodal LLMs to jointly perform captioning, VQA, and grounding. It achieves 82% accuracy on captioning and 83.32% on VQA, with 90.94%, 52.04%, and 92.06% for binary, numeric, and semantic question types respectively. For grounding, it attains 64.94% accuracy. By combining diverse VLMs with our custom RemoteSAM-SAM3 models through ensemble majority voting, the system delivers more accurate and consistent remote-sensing understanding than task-specific approaches.