arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RRS-10K:用于罕见遥感图像解释的多任务视觉语言模型基准测试

RRS-10K: A Multitask Vision-Language Model Benchmark for Rare Remote Sensing Image Interpretation

Yuqiao Lai, Jiancheng Qi, Fei Wang, Yuxin Liu, Kun Li, Ye Chen, Yan Gao, Yanyan Wei

arXiv 2607.24810首次发表:更新:

发表机构

Laboratory of Intelligent Language Processing, National University of Defense Technology; Hefei University of Technology; Institute of Artificial Intelligence, Hefei Comprehensive National Science Center; United Arab Emirates University(国防科技大学智能语言处理实验室; 合肥工业大学; 合肥综合性国家科学中心人工智能研究院; 阿联酋大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视觉语言模型在罕见遥感图像解释能力不足的问题,提出RRS-10K基准测试,含军事遥感图像及问答对,构建时引入干扰项过滤策略,评估多个模型,揭示其性能弱点,为开发更可靠的遥感视觉语言模型提供指导。

AI 中文摘要

视觉语言模型在一般遥感任务中表现出色,但对罕见场景的能力了解不足,因为现有基准测试多为常见城乡图像。为此提出RRS-10K基准测试,包含10738张军事相关遥感图像及多格式问答对,按三个能力维度等组织。构建基准时引入基于相似性的干扰项过滤策略。评估52个模型发现当前视觉语言模型在罕见遥感图像解释上零样本性能一般,RRS-10K能系统分析长尾遥感解释中的失败模式并指导模型开发。

英文摘要

Vision-language models (VLMs) have achieved strong performance on general remote sensing tasks. However, their capability for rare scenes remains insufficiently understood, because existing benchmarks are dominated by common urban and rural imagery. To address this gap, we present RRS-10K, a benchmark for rare remote sensing image interpretation. RRS-10K contains 10,738 military-related remote sensing images and corresponding multiple format question-answer pairs for comprehensive evaluation. All of the images are collected from first-hand sources and organized into three capability dimensions, six sub-dimensions, and 20 leaf tasks, covering perception, reasoning, and robustness. To improve the quality of multiple-choice questions, we introduce a similarity-based distractor filtering strategy (SDFS) during benchmark construction. We further evaluate 52 representative models and show that current VLMs achieve only moderate zero-shot performance on rare remote sensing image interpretation, with clear weaknesses in visual grounding, referring segmentation, and complex semantic reasoning tasks. RRS-10K enables systematic analysis of failure modes in long-tail remote sensing interpretation and provides guidance for developing more reliable remote sensing VLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑