arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于遥感图像理解的多模态大语言模型:领域特定还是通用?

Multimodal Large Language Models for Remote Sensing Image Understanding: Domain-Specific or General-Purpose?

Qiwei Ma, Chunping Qiu, Xinjun Cheng, Xiaoyu Zhang, Puhong Duan, Ke Yang, Xudong Kang, Shutao Li

arXiv 2607.20284首次发表:更新:

发表机构

School of Artificial Intelligence and Robotics, Hunan University; Intelligent Game and Decision Lab (IGDL); Yuelushan Center for Industrial Innovation(湖南大学人工智能与机器人学院; 智能游戏与决策实验室; 岳麓山工业创新中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对遥感图像理解的多模态大语言模型展开研究,通过系统调查和评估,比较其与通用模型在不同任务上的表现,发现当前模型存在局限,进而给出未来方向,为开发相关模型提供系统参考。

AI 中文摘要

多模态大语言模型(MLLMs)的快速发展为遥感图像场景理解(RSISU)带来了灵活范式,实现与遥感图像的自然语言交互。然而,对现有遥感MLLMs(RS - MLLMs)的能力边界、跨任务泛化和特定任务局限性仍缺乏系统理解。本文对用于RSISU的MLLMs进行系统调查和诊断评估。回顾RS - MLLMs技术演变,关注模型设计等。比较RS - MLLMs与通用计算机视觉MLLMs(CV - MLLMs)在不同RSISU任务和基准上的表现。发现RS - MLLMs在特定领域有竞争力,通用CV - MLLMs在一些任务上可匹配甚至超越。当前MLLMs在空间和关系推理等方面也有局限。基于此给出未来方向,为开发用于RSISU的强大、通用且实用的MLLMs提供系统参考。

英文摘要

The rapid development of multimodal large language models (MLLMs) has introduced a flexible paradigm for remote sensing image scene understanding (RSISU), enabling natural-language interaction with remote sensing imagery. However, a systematic understanding of the capability boundaries, cross-task generalization, and task-specific limitations of existing remote sensing MLLMs (RS-MLLMs) is still lacking. This paper presents a systematic survey and diagnostic evaluation of MLLMs for RSISU. We review the technical evolution of RS-MLLMs, focusing on model design, multimodal learning, training data, and downstream capabilities. We further compare RS-MLLMs with general-purpose computer vision MLLMs (CV-MLLMs) across diverse RSISU tasks and benchmarks. RS-MLLMs remain competitive in domain-specific settings, particularly remote sensing visual grounding and high-resolution visual question answering. More notably, general-purpose CV-MLLMs can match or even outperform these specialized models on several RSISU tasks without remote sensing-specific fine-tuning. These findings demonstrate the strong transferability of general-purpose CV-MLLMs and show that current RS-MLLMs do not consistently outperform them across diverse RSISU tasks. Current MLLMs also face limitations in spatial and relational reasoning, fine-grained visual understanding, instruction diversity, and generalization across heterogeneous task formats. Based on these findings, we outline future directions toward reliable evaluation, multimodal and high-resolution reasoning, efficient deployment, and tool-augmented remote sensing agents. This survey provides a systematic reference for developing robust, generalizable, and practical MLLMs for RSISU.

Comments27 pages, 11 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑